A video from a photo and a voice
One photo and one audio clip are enough. There is no studio booking, no on-camera talent, and no motion-capture rig between an idea and a finished presenter video.
OmniHuman is ByteDance's audio-driven human video model. It takes a single photo of a person and a speech or audio track, and generates a full-body video of that subject talking, gesturing, and moving with natural, lip-synced motion.
In Phygital+ it runs as a node with a required Character Image and Audio input, plus an optional Text Prompt for extra movement guidance and an optional Mask for selecting one subject in a busier image. Fast Mode and a fixed Seed are exposed for speed and reproducibility.
Avatar presenters, narrated explainers, product announcements, and character-based social clips
Content, education, and marketing teams that need a face and a voice without filming
Turn a single photo and an audio track into a full-body talking video with natural gestures and lip-sync.
One photo and one audio clip are enough. There is no studio booking, no on-camera talent, and no motion-capture rig between an idea and a finished presenter video.
The model was trained with a mixed multi-signal strategy across audio, pose, and video, which is why it holds up on portrait, half-body, and full-body framing instead of only animating a face.
Lip movement tracks the connected audio closely, so narration, dialogue, or a translated voice-over reads as a real performance rather than a mismatched overlay.
OmniHuman animates stylized illustrations and humanoid characters as well as real photographs, so a brand mascot or 2D character can deliver the same line a live presenter would.
Mask lets you pick one subject out of a busier photo, and Text Prompt adds a gesture or action on top of what the audio already drives, without needing a second take.
OmniHuman fits wherever a message needs a face and a voice, without booking a studio or an on-camera presenter.
Connect a character photo and an audio track, add optional guidance, and generate.
One subscription across 30+ AI models, with no per-tool credit balances or separate signups. Credit cost per generation is shown live in the node before you run it.
500 weekly credits to test the model and see the output quality. No credit card required.
45,000 monthly credits with cheaper per-credit pricing, commercial rights, and video download.
90,000 monthly credits, up to 10 seats, shared workspace, centralized billing, and priority support.
210,000 monthly credits, API and integration support, unlimited seats, and dedicated management.
Full specs for OmniHuman in Phygital+.
How OmniHuman compares with other AI video and avatar tools in the Phygital+ catalog.
Chain scripting, voice, and avatar animation into one repeatable workflow.
Design or generate the character portrait with an image node, then bring it straight into OmniHuman to animate.
Write the script and generate a voice-over in Sound Creation or Voice Clone, then connect that audio to OmniHuman.
Take the finished clip into a video upscaler to reach delivery resolution for social or broadcast.
Draft the script with a text node before turning it into speech and then into an animated presenter.
Everything you need to know about OmniHuman in Phygital+.
OmniHuman is ByteDance's human video generation model, first published in February 2025 as OmniHuman-1. It turns a single image of a person and a speech or audio track into a realistic video where the subject talks, gestures, and moves in sync with the sound, without recording any real footage.
OmniHuman-1 uses a Diffusion Transformer backbone trained with an omni-conditions strategy: instead of learning only from perfectly clean data, it mixes audio, pose, and video motion signals during training. That lets it generalize to portrait, half-body, and full-body shots, and to stylized or non-human characters, more reliably than earlier face-only or upper-body-only avatar models.
Yes. The model was trained on humanoid and stylized characters as well as real people, and can animate 2D cartoon figures or anthropomorphic subjects, not only photographic portraits.
In Phygital+, connect a Character Image and an Audio track; both are required. Text Prompt is optional and steers movement, action, or posture. Mask is optional and restricts animation to a specific person or region when the image contains more than one subject.
Under 35 seconds. The node accepts audio up to 60 seconds long in the field definition, but keeping the clip under 35 seconds is the supported range for a single generation. For a longer script, split it into several generations.
It is optional and works alongside the required image and audio, not instead of them. Use it to describe body movement, a gesture, or a change in posture; the audio still drives timing and lip-sync.
Enable Fast Mode when turnaround speed matters more than preserving every motion detail; it speeds up generation at some cost to effect quality. Leave it off for the highest-fidelity result, and fix a Seed if you need to reproduce a specific take.
Usage is credit-based and included in every plan, from Free up to Enterprise. The credit cost per generation is shown in the node before you run it, so there is no separate API account to set up.
Outputs created on any paid plan come with commercial-use rights. Make sure you have the rights to the photo and voice you animate, especially for any real person's likeness.
Browse the full catalog of 30+ AI models available in Phygital+.
Join 100+ teams using Phygital+ – every model in one workspace
Try Phygital+ free