OmniHuman 1.5 guide: how to turn one image and audio into a talking avatar video

OmniHuman starts with two familiar assets: one image and one audio track. The image defines who appears on screen. The audio provides the words, rhythm, pauses, and emotional cues that drive the performance.

ByteDance designed OmniHuman 1.5 as an audio-driven avatar model that can animate more than a face. It can produce visible gestures, posture changes, full-body movement, and camera decisions that respond to the content of the audio. An optional text prompt adds direction for expression, action, and framing.

This makes the model useful for presenters, explainers, product announcements, translated videos, musical performances, fictional characters, and social content built from an approved visual identity.

In this guide, we will cover the differences between OmniHuman-1 and 1.5, how to choose an image and audio track, how optional prompts work, practical examples, and how to connect the model to a repeatable Phygital+ pipeline.

OmniHuman-1 compared with OmniHuman 1.5

The original OmniHuman-1 research introduced a Diffusion Transformer framework trained with mixed motion conditions. Its main achievement was flexible human animation across close portraits, half-body shots, full-body shots, talking, singing, difficult poses, and different image styles.

OmniHuman 1.5 adds stronger semantic control. The model uses the meaning, rhythm, prosody, and emotion of the audio together with image and text guidance to create a more directed performance.

VersionMain strengthBest use
OmniHuman-1Flexible audio-driven human animation across varied framing and stylesTalking and singing portraits, half-body and full-body animation
OmniHuman 1.5Audio semantics, optional text control, richer gestures, emotion, and camera movementDirected presenters, performances, character videos, and campaign content
OmniHuman 1.5 in Phygital+Image and audio inputs with optional prompt, mask, seed, and Fast ModeRepeatable node-based avatar workflows connected to image, voice, and video tools
Current model rule: choose OmniHuman 1.5 when the performance needs to react to what is being said, and use the optional prompt when you need a specific gesture, posture, camera move, or visual event.

What OmniHuman 1.5 is good at

Audio-driven lip-sync and expression

The audio is a driving condition, rather than a soundtrack added after animation. Mouth movement, facial expression, pauses, and body language are generated around the connected voice or music.

Portrait, half-body, and full-body framing

A close portrait gives the model more facial detail. Half-body framing leaves room for hand and upper-body gestures. Full-body images support larger movement but require a clear silhouette and enough space around the subject.

Talking, singing, and rhythmic performance

ByteDance’s official materials show dialogue, emotional acting, and musical performance. For singing, a clean vocal track and an image that already fits the intended performance usually give the model a clearer foundation.

Real people, illustrations, and characters

The source image can be photographic or stylized. Cartoon characters, mascots, game characters, and some animal subjects can work when the face and body remain readable.

Text-directed action and camera

The optional text prompt can guide movement, emotion, posture, camera motion, environmental interaction, or an element that does not appear in the source image.

How to prepare the image and audio

Choose a readable source image

  • Use a sharp image with a clearly visible face.
  • Keep hands and limbs visible if the prompt asks for gestures.
  • Avoid heavy foreground obstruction across the mouth or body.
  • Leave space around the subject for camera movement.
  • Use an image whose lighting and setting already fit the final video.

Prepare clean audio

  • Use a clear voice with minimal background noise.
  • Remove long silence from the beginning and end.
  • Keep pacing natural; very fast speech makes lip-sync and gesture timing harder.
  • Use one deliberate emotional direction per clip.
  • Split long scripts into scenes that can be reviewed separately.

Official API implementations commonly expose 720p for audio up to 60 seconds and 1080p for audio up to 30 seconds. Limits inside a particular interface can be stricter, so the active Phygital+ node remains the final source for the allowed file and duration settings.

Use a mask for crowded images

When the image contains several people, the optional Mask input in Phygital+ can identify the subject that should animate. Multi-person performance is a model-level capability, but it requires separate subject and audio routing. A single mixed audio track is not a reliable substitute for speaker assignment.

How to prompt OmniHuman 1.5

The image defines the visual identity and the audio drives timing. The optional text prompt should focus on the performance decisions that are not already obvious from those inputs.

Prompt layerWhat to describeExample
ExpressionEmotional state and how it changesBegins calmly, then smiles with quiet confidence
GestureHands, posture, or body movementUses small open-hand gestures while explaining
CameraFraming and movementSlow push-in from a medium shot to a close portrait
InteractionHow the subject uses the environmentTurns toward the product and points to its label
StyleOverall performance and visual treatmentNatural UGC delivery with subtle handheld movement
ContinuityDetails that should remain stableKeep the face, outfit, background, and lighting unchanged
Prompt formula: [Framing and camera]. The subject delivers the connected audio with [expression and energy]. They use [specific gestures or action]. Keep [identity and scene details] unchanged.

Short prompts are often enough. The audio already carries timing and emotion, so several contradictory performance directions can make the result less coherent.

OmniHuman 1.5 prompt examples

Product presenter

Medium shot. The presenter speaks with calm confidence, uses small open-hand gestures, then turns slightly toward the product and points to its label. Slow camera push-in. Preserve the face, outfit, product, and studio lighting.

UGC review

Natural vertical UGC delivery. The creator speaks directly to the camera, smiles briefly after the opening line, and lifts the product into view at the key benefit. Subtle handheld camera feel and realistic pauses.

Course teacher

Half-body educational presenter. Clear, measured delivery with restrained gestures. The teacher turns toward the diagram on the right during the explanation, then returns attention to the camera.

Company update

Professional spokesperson in a bright office. Upright posture, steady eye contact, and concise hand gestures timed to the main points. Fixed medium camera and clean corporate presentation.

Founder message

Intimate founder message filmed at a desk. Warm, direct delivery with natural breathing and subtle head movement. Slow push-in during the final sentence. Keep the workspace and daylight stable.

News presenter

Centered news-style presenter. Neutral expression at the opening, more urgency in the middle, then a calm conclusion. Minimal gestures and stable eye-level camera.

Singer

Emotional live vocal performance. The singer follows the rhythm with natural upper-body movement, closes their eyes during the sustained note, then looks back toward the camera. Gentle cinematic arc movement.

Podcast clip

Seated podcast guest in a three-quarter view. Conversational delivery with occasional eyebrow movement and one small hand gesture during the strongest point. Static camera and realistic studio ambience.

Historical portrait

Animate the historical portrait as a restrained first-person monologue. Formal posture, subtle facial motion, and period-appropriate composure. Keep the painted texture and original background intact.

Brand mascot

The illustrated mascot delivers the audio with upbeat energy, nods at the main claim, and points toward the logo at the end. Preserve the original illustration style, colors, proportions, and face.

Fantasy character

Cinematic fantasy monologue. The character begins in profile, turns toward the camera during the warning, and tightens one hand around the pendant. Slow orbiting camera and flickering firelight.

Pet character

Playful talking pet video. The dog tilts its head between phrases, looks toward the treat on the table, and reacts with excited but physically believable movement. Keep the fur pattern and room unchanged.

Full-body host

Full-body presenter in a gallery. The host walks slowly toward the first display while speaking, stops beside it, and uses one clear presenting gesture. Smooth side tracking camera.

E-commerce demo

The presenter demonstrates the bag while delivering the connected voice-over. They open the main compartment, show the interior, and return the product to a centered position. Preserve the bag shape, material, and logo.

Emotional actor

Close cinematic performance. The actor begins guarded and quiet, becomes visibly emotional during the middle line, then regains composure before the final sentence. Very slow push-in and soft window light.

Building an OmniHuman workflow in Phygital+

OmniHuman becomes more useful when the image and audio are prepared inside the same canvas. A text model can write and shorten the script. An image model can create the presenter, mascot, or campaign character. Voice Creation or Voice Clone can produce the audio. OmniHuman combines the approved image and voice, and a video upscaler can prepare the chosen result for delivery.

  1. Draft the script and split it into reviewable scenes.
  2. Create or select the character image.
  3. Generate or upload the voice track.
  4. Connect image and audio to OmniHuman.
  5. Add an optional motion prompt and mask if needed.
  6. Generate several takes with controlled changes.
  7. Select the strongest performance and continue to upscale or editing.

Keeping these steps visible makes revisions easier. A new voice, language, gesture, or presenter can branch from the approved materials without rebuilding the project.

OmniHuman pricing in Phygital+

OmniHuman uses Phygital credits. The exact generation cost is shown inside the node before the run, so you can review the current cost without creating a separate BytePlus account.

Test the image, audio quality, framing, and optional prompt on a short segment before generating every scene in a long script.

Common OmniHuman problems

The lip-sync looks weak

Use clean speech, remove music that competes with the voice, and avoid very fast delivery. A clearly visible mouth in the source image also helps.

The hands or body move unnaturally

Use a source image with readable anatomy and ask for one restrained gesture at a time.

The identity changes

Start from a sharp image and avoid prompts that request major changes to face, hair, wardrobe, or viewing angle.

The result feels static

Use half-body or full-body framing and describe one camera move or environmental interaction.

The wrong person moves

Supply a mask that isolates the intended subject, or crop the source image to remove ambiguity.

The performance does not match the message

Describe the emotional arc in plain language and use audio with clear prosody. The model responds to the meaning and delivery of the voice.

Current limitations

Audio-driven generation can struggle with extreme head turns, hidden mouths, crowded compositions, rapid choreography, intricate hand-object contact, and contradictory text guidance. Multi-person scenes need explicit speaker routing; one audio track cannot reliably assign different lines to several people.

Rights and consent also matter. Use images and voices that you own or are permitted to animate, especially when the subject is a real person.

FAQ

What is OmniHuman 1.5?

OmniHuman 1.5 is ByteDance’s audio-driven avatar video model. It uses an image, audio, and optional text guidance to generate a synchronized character performance.

What inputs does OmniHuman need?

The Phygital+ node requires a character image and an audio track. Text Prompt and Mask are optional controls.

Can OmniHuman animate a full body?

Yes. The model supports portrait, half-body, and full-body source images. Clear anatomy and an unobstructed silhouette improve larger movements.

Can OmniHuman animate illustrations or mascots?

Yes. Official materials include stylized, cartoon, non-human, and animal subjects as well as photographic people.

Does OmniHuman support 1080p?

Yes. BytePlus describes native 1080p output. Common API limits allow up to 30 seconds at 1080p and up to 60 seconds at 720p, while interface limits can vary.

What does the optional text prompt control?

It can guide expression, gesture, posture, action, camera movement, style, and interaction with the scene.

Can OmniHuman create singing videos?

Yes. The model supports rhythmic performance and singing driven by an audio track.

Can several people speak in one video?

OmniHuman 1.5 supports multi-person scenarios at the model level, but each speaker needs explicit subject and audio routing. Use a mask or separate generations when the interface exposes one primary speaker.

How do I keep the same character across several clips?

Reuse the same approved image, keep wardrobe and scene descriptions stable, and change one performance variable at a time.

Why use OmniHuman in Phygital+?

Phygital+ connects scripting, image generation, voice creation, OmniHuman animation, branching, and upscaling on one visual canvas.

Explore more