Image-to-video Audio-driven ByteDance Commercial use
OmniHuman

OmniHuman
– Talking avatar video
model in Phygital+

OmniHuman is ByteDance's audio-driven human video model. It takes a single photo of a person and a speech or audio track, and generates a full-body video of that subject talking, gesturing, and moving with natural, lip-synced motion.

In Phygital+ it runs as a node with a required Character Image and Audio input, plus an optional Text Prompt for extra movement guidance and an optional Mask for selecting one subject in a busier image. Fast Mode and a fixed Seed are exposed for speed and reproducibility.

OmniHuman

At a glance OmniHuman

Built for

Avatar presenters, narrated explainers, product announcements, and character-based social clips

Content, education, and marketing teams that need a face and a voice without filming

Vendor
ByteDance
Released
February 2025 (OmniHuman-1)
Category
AI video — audio-driven human animation
Modalities
Image + audio (+ optional text, mask) → video
Audio length
Under 35 seconds per generation
Subjects
Real people, humanoid, and stylized characters
Commercial use
Yes
Access
All Phygital+ plans

Why use OmniHuman

Turn a single photo and an audio track into a full-body talking video with natural gestures and lip-sync.

1

A video from a photo and a voice

One photo and one audio clip are enough. There is no studio booking, no on-camera talent, and no motion-capture rig between an idea and a finished presenter video.

2

Full-body motion, not just a face

The model was trained with a mixed multi-signal strategy across audio, pose, and video, which is why it holds up on portrait, half-body, and full-body framing instead of only animating a face.

3

Lip-sync that holds up

Lip movement tracks the connected audio closely, so narration, dialogue, or a translated voice-over reads as a real performance rather than a mismatched overlay.

4

Works beyond real people

OmniHuman animates stylized illustrations and humanoid characters as well as real photographs, so a brand mascot or 2D character can deliver the same line a live presenter would.

5

Control without re-shooting

Mask lets you pick one subject out of a busier photo, and Text Prompt adds a gesture or action on top of what the audio already drives, without needing a second take.

Built for content, education, and marketing teams

OmniHuman fits wherever a message needs a face and a voice, without booking a studio or an on-camera presenter.

OmniHuman
edu
OmniHuman
OmniHuman
edu
OmniHuman

How to use OmniHuman

Connect a character photo and an audio track, add optional guidance, and generate.

OmniHuman
Step 1

Connect image and audio

Add an OmniHuman node to the canvas and connect a clear photo of the person or character to Character Image, and a speech or audio clip under 35 seconds to Audio. Both inputs are required.

OmniHuman
Step 2

Add optional guidance

Add Text Prompt if you want to guide gestures, posture, or a specific action, and Mask if the image contains more than one person and only one should animate.

OmniHuman
Step 3

Set parameters, generate & export

Fix a Seed to reproduce a result, or enable Fast Mode when speed matters more than preserving every motion detail. Run the node and review the video output, then chain it into an upscaler or export it.

Generated with OmniHuman

OmniHuman
OmniHuman
OmniHuman
OmniHuman
OmniHuman
OmniHuman
OmniHuman
OmniHuman
Start Generating!

Pricing and access

One subscription across 30+ AI models, with no per-tool credit balances or separate signups. Credit cost per generation is shown live in the node before you run it.

Free

Try OmniHuman free

500 weekly credits to test the model and see the output quality. No credit card required.

  • 500 weekly credits
  • Access to 30+ models
  • Personal use
Join as Free
Pro

Professional access

45,000 monthly credits with cheaper per-credit pricing, commercial rights, and video download.

  • 45,000 monthly credits
  • Commercial use license
  • 15% cheaper credits
Join as Pro
Team

Team collaboration

90,000 monthly credits, up to 10 seats, shared workspace, centralized billing, and priority support.

  • 90,000 monthly credits
  • Up to 10 seats
  • Centralized billing
Join as Team
Enterprise

Enterprise scale

210,000 monthly credits, API and integration support, unlimited seats, and dedicated management.

  • API & integration support
  • Unlimited seats
  • Dedicated manager
Join as Enterprise

Technical Specifications

Full specs for OmniHuman in Phygital+.

Model
OmniHuman-1 / OmniHuman 1.5
Released
February 2025 (research paper + model)
Required inputs
Character Image, Audio
Output
Video with synced lip movement and gestures
Seed
-1 to 2147483647, random by default
Subject support
Human, humanoid, and stylized/cartoon characters
Vendor
ByteDance
Architecture
Diffusion Transformer (DiT) with omni-conditions training
Optional inputs
Text Prompt, Mask
Audio duration
Under 35 seconds recommended, 60s field max
Fast Mode
Optional toggle, trades effect quality for speed

OmniHuman vs Other AI models

How OmniHuman compares with other AI video and avatar tools in the Phygital+ catalog.

Model

OmniHuman

One image plus one audio track, turned into a full-body talking video

Kling Omni

Character-consistent scenes with multi-shot storyboards, no dedicated avatar input

HeyGen

Studio-recorded presenters with pre-built avatar library

Wan Video

Text-to-video and image-to-video without a dedicated talking-avatar mode

Best for Any photo becomes a talking, gesturing presenter Directed multi-shot sequences with named elements Pre-built presenter avatars, script-to-video General text-to-video and image-to-video
Input One image + one audio track (speech-driven) Text prompt, reference images Text script, avatar selection Text prompt, start/end frame
Motion Full-body gestures, expression, and precise lip-sync Not audio-driven by default Stock and custom avatars, not your own photo animated freely No lip-sync focus
Duration Up to 35 seconds of audio per generation Multiple minutes per shot Minutes Up to 20 seconds

Use OmniHuman Alongside other AI models In Phygital+

Chain scripting, voice, and avatar animation into one repeatable workflow.

OmniHuman

Chain with an image model

Design or generate the character portrait with an image node, then bring it straight into OmniHuman to animate.

OmniHuman

Chain with voice generation

Write the script and generate a voice-over in Sound Creation or Voice Clone, then connect that audio to OmniHuman.

OmniHuman

Chain with a video upscaler

Take the finished clip into a video upscaler to reach delivery resolution for social or broadcast.

OmniHuman

Chain with a text model

Draft the script with a text node before turning it into speech and then into an animated presenter.

Start Generating!

Everything you need to know about OmniHuman in Phygital+.

FAQ about OmniHuman

OmniHuman
What is OmniHuman?

OmniHuman is ByteDance's human video generation model, first published in February 2025 as OmniHuman-1. It turns a single image of a person and a speech or audio track into a realistic video where the subject talks, gestures, and moves in sync with the sound, without recording any real footage.

What makes OmniHuman different from older talking-head tools?

OmniHuman-1 uses a Diffusion Transformer backbone trained with an omni-conditions strategy: instead of learning only from perfectly clean data, it mixes audio, pose, and video motion signals during training. That lets it generalize to portrait, half-body, and full-body shots, and to stylized or non-human characters, more reliably than earlier face-only or upper-body-only avatar models.

Does it work on cartoon or stylized characters?

Yes. The model was trained on humanoid and stylized characters as well as real people, and can animate 2D cartoon figures or anthropomorphic subjects, not only photographic portraits.

What do I connect to the OmniHuman node?

In Phygital+, connect a Character Image and an Audio track; both are required. Text Prompt is optional and steers movement, action, or posture. Mask is optional and restricts animation to a specific person or region when the image contains more than one subject.

How long can the audio be?

Under 35 seconds. The node accepts audio up to 60 seconds long in the field definition, but keeping the clip under 35 seconds is the supported range for a single generation. For a longer script, split it into several generations.

What does the Text Prompt field do?

It is optional and works alongside the required image and audio, not instead of them. Use it to describe body movement, a gesture, or a change in posture; the audio still drives timing and lip-sync.

What does Fast Mode change?

Enable Fast Mode when turnaround speed matters more than preserving every motion detail; it speeds up generation at some cost to effect quality. Leave it off for the highest-fidelity result, and fix a Seed if you need to reproduce a specific take.

How much does OmniHuman cost in Phygital+?

Usage is credit-based and included in every plan, from Free up to Enterprise. The credit cost per generation is shown in the node before you run it, so there is no separate API account to set up.

Can I use the videos commercially?

Outputs created on any paid plan come with commercial-use rights. Make sure you have the rights to the photo and voice you animate, especially for any real person's likeness.

Explore more AI models In Phygital+

Browse the full catalog of 30+ AI models available in Phygital+.

Dialogue Creation ElevenLabs v3 FLUX Gemini Omni GPT-Image-2 Hailuo 3.0 Ideogram 3 Kling 3.0 Kling Omni Kling Omni Image Krea AI LTX Video Luma Magnific Upscale Midjourney Nano Banana Pro OmniHuman Qwen Image Recraft V4 Reve Runway 4.5 Seedance 2.5 Seedream 5 Sound Creation Topaz Upscale Upscale Video Veo 3.1 Voice Clone Voice Creation WAN 2.5 Wan Video 2.1

Ready to work Faster with OmniHuman?

Join 100+ teams using Phygital+ – every model in one workspace

Try Phygital+ free
bg