Hedra
  • Developers
  • Studio
  • Enterprise
  • Blog
Log inSign Up
Open Hedra
Your account
    Explore
  • Home
  • Developers
  • Studio
  • Enterprise
  • Blog
    Log inSign Up
    Open Hedra
A medium close-up video frame of a person with a realistic falcon head sitting at a podcast desk. The figure wears a blue button-down shirt and sits before a black microphone on a boom arm. The background is a studio with acoustic panels and warm orange lights, generated with Omnihuman 1.5 at 1920x1088 resolution.

Omnihuman 1.5

Creates vivid, emotional character videos driven entirely by your audio.

Get an API keyOpen Creative Studio
All models

Overview

Developed by ByteDance, Omnihuman 1.5 is a film-grade video model that transforms a single image and an audio track into a highly expressive digital human. The model uses a dual-system cognitive architecture to generate realistic lip-syncing, context-aware gestures, and continuous camera movement. It is especially good for producing lifelike avatars, virtual actors, and multi-character interactions for storytelling, e-commerce, and marketing.

Specifications · Text + image + audio to video

Input mode
Text + image + audio to video
Accepts
start frame, audio (required)
Aspect ratios
16:9
Resolutions
720p, 1080p
Native audio
Audio-driven

Pricing

ByteDance
720p/1080p16¢/second

Build with this model: the Omnihuman 1.5 API on the Hedra Developer Platform.

Omnihuman 1.5 API →

Real output · generated with Omnihuman 1.5

Best of Omnihuman 1.5

A medium close-up video frame of a person with a realistic falcon head sitting at a podcast desk. The figure wears a blue button-down shirt and sits before a black microphone on a boom arm. The background is a studio with acoustic panels and warm orange lights, generated with Omnihuman 1.5 at 1920x1088 resolution.Falcon Podcaster Speaking Into Microphone — Omnihuman 1.5

Prompting

Prompt tips

  • Write like a screenplay: Structure your prompts sequentially: [Camera movement] + [Emotion] + [Speaking state] + [Specific actions] (e.g., "Camera pushes in. She speaks thoughtfully, pausing mid-sentence with a slight smile").
  • Focus on action verbs: Do not describe the character's physical appearance, as the model relies on the source image. Focus entirely on movement, behavior, and camera direction.
  • Draft in 720p: Use the 720p resolution to quickly iterate on your prompt's timing and camera moves, then switch to 1080p for the final high-quality render.
  • Use explicit speaking cues: If the lip-sync feels slightly off, add explicit verbs like "speaking," "singing," or "listening quietly" to firmly guide the model's facial generation.
  • Generate source images first: Use a high-fidelity image model like Dreamina 3.1 or Flux 1.1 Pro to create a clean, well-lit portrait before animating it here.

About the model

Questions, answered

What is Omnihuman 1.5 best used for?

Omnihuman 1.5 is a video generation model by ByteDance designed to create realistic, lip-synced digital avatars from a single portrait image and an audio track. It excels at generating expressive performances where the character's facial expressions, body gestures, and head movements match the rhythm and emotional tone of the speech. It is highly effective for creating virtual presenters, personalized video messages, and cinematic talking-head videos, and it can even animate stylized non-human characters like anime figures or pets.

What makes Omnihuman 1.5 different from other lip-sync models?

It utilizes a "cognitive simulation" architecture inspired by human psychology. Instead of mechanically matching lips to audio waveforms, a multimodal large language model analyzes the semantic meaning of the audio to plan appropriate emotional reactions and gestures. Then, a diffusion transformer renders the physical movements. This dual-system approach allows the avatar to appear as if it is thinking and reacting naturally to the context of the speech, reducing the stiff, robotic feel common in older avatar models.

Who developed Omnihuman 1.5 and when was it released?

Omnihuman 1.5 was developed by ByteDance's Intelligent Creation team. The model's research paper and core architecture were unveiled on August 26, 2025. It serves as a major architectural upgrade over the original OmniHuman-1, adding text-prompt guidance, unconstrained camera movement, and cognitive reasoning. ByteDance is also the developer behind other notable generative models, including the image generator Dreamina 3.1 and the multimodal video model Seedance 2.0.

How can I create a multi-character conversation using Omnihuman 1.5?

According to the official BytePlus documentation, Omnihuman 1.5 cannot use a single audio file to drive a back-and-forth conversation between multiple characters simultaneously. To achieve this, you must use subject detection to isolate each character with a mask, generate individual speaking clips for each person using their specific audio segments, and then stitch the resulting videos together in a video editor.

Can I guide the avatar's performance beyond just providing audio?

Yes. A key upgrade in version 1.5 is the addition of text prompt support. You can provide a text prompt alongside your image and audio to explicitly direct the character's emotional state, body language, and even camera movements, such as zooms or pans. This gives you directorial control over the final performance rather than relying entirely on the AI's automatic interpretation of the audio track.

Same API · one key

Similar models

A close-up shot of a battle-worn male warrior with curly dark hair and light facial scars, wearing dark leather armor. He has his eyes closed in exhaustion, with embers floating around him in a smoky battlefield setting. This 1080x1080 video was generated using MiniMax Hailuo 2.3 Fast Pro on Hedra.MiniMax Hailuo 2.3 Fast ProMiniMaxA close-up shot of a dark-haired man in a heavy black fur-trimmed cloak, standing against a backdrop of blurry, snow-covered mountains. Generated using MiniMax Hailuo 2.3 Fast Standard at a resolution of 1024x768, this video frame depicts a cinematic fantasy character with a serious expression.MiniMax Hailuo 2.3 Fast StandardMiniMaxA medium close-up shot of a smiling young woman with long wavy brown hair wearing a cream-colored knit sweater. She is captured in a bedroom setting, looking directly at the camera in a vlog style. This 1280x720 video was generated using Grok Video on Hedra.Grok VideoxAIA wide-angle shot of ocean waves crashing onto a sandy beach at sunset, with distant cliffs on the horizon. The orange and pink sky is reflected on the wet shoreline. This 1924x1076 video frame was generated using the Kling 2.1 Master model.Kling 2.1 MasterKlingFirst frame of a 10-second video generated using Kling 2.5 Turbo at 1924x1076 resolution. An elderly, silver-haired scholar wearing a dark red blazer leans over a wooden desk covered in old maps and brass compasses. In the background, a globe and a faded map of the world are visible under warm, dusty light.Kling 2.5 TurboKlingA widescreen 1080x720 seated talking-animal video of a large bear in a work shirt, delivering a gruff, matter-of-fact roofing estimate to camera. Generated from a still image and an audio track using Hedra Character 3.Hedra Character 3Hedra

What will you create?

Get an API keyOpen Creative Studio
Hedra
Hedra

Product

Developer PlatformStudioEnterpriseSovereignPricing

Resources

Agent documentationDeveloper documentationBlogUse CasesModelsFeedbackChangelogStatus

Company

AboutCareersContactBuilder programSupportAlternatives

Legal

Privacy PolicyTerms of useAcceptable useCookie PolicyBiometric data policy
LinkedinInstagramDiscord
support@hedra.comHedra 2026 — All rights reserved