WAN 2.5 combines short-form video generation with synchronized sound. A prompt can define the scene, movement, camera, dialogue, ambience, and music, while an optional audio file can drive timing and lip movement.
The model supports text-to-video and first-frame image-to-video. Official Alibaba Cloud documentation lists 5- and 10-second clips, 480p, 720p, and 1080p output, 30 frames per second, and MP4 delivery.
This guide focuses on using that compact format well: writing one clear scene, planning the sound with the image, choosing between text and image input, and building controlled branches in Phygital+.
WAN versions compared
The Wan family changes quickly. WAN 2.5 remains useful because it combines multiple resolution tiers with audio generation and custom audio synchronization. WAN 2.6 is newer and adds longer duration ranges and multi-shot capabilities in supported modes.
| Version | Main strength | Best use |
|---|---|---|
| WAN 2.1 | Earlier widely available text-to-video and image-to-video family | Existing and self-hosted workflows tied to 2.1 variants |
| WAN 2.2 | Improved visual generation with fixed short clips in documented modes | Short visual-first generation |
| WAN 2.5 | 5/10-second video, 480p–1080p, native or uploaded audio | Short ads, dialogue, music, product clips, and social video with sound |
| WAN 2.6 | Newer family with longer duration and multi-shot support in selected modes | Sequences that need more time or multiple connected shots |
WAN 2.5 generation modes
Text-to-video
Text-to-video starts from the brief. Describe one subject, one main action, the environment, camera, lighting, and audio. This mode works well for concept exploration when no exact starting image is required.
Image-to-video
Image-to-video uses an image as the first frame. The prompt should describe motion and sound instead of repeatedly describing what is already visible. This is the stronger choice for product shape, character identity, composition, or campaign art direction.
Resolution and duration
- Durations: 5 or 10 seconds
- Resolution tiers: 480p, 720p, and 1080p
- Frame rate: 30 fps
- Format: MP4 using H.264 encoding
- Common aspect ratios include 16:9, 9:16, 1:1, 4:3, and 3:4 depending on the resolution tier
Use the lower resolution for inexpensive direction tests when the interface exposes it, then move the approved prompt or starting image to the delivery resolution.
Native audio and custom audio
Automatic dubbing
If no audio file is supplied, WAN 2.5 can generate sound from the prompt. This can include dialogue, environmental effects, music, and ambience. State exactly what should be audible and remove sounds that do not serve the scene.
Uploaded audio
An uploaded WAV or MP3 file can guide the result. Alibaba documents custom audio from 3 to 30 seconds, although the generated video remains 5 or 10 seconds. Audio beyond the selected video duration is cut; if it is shorter, the remaining section can be silent.
Dialogue prompting
Place spoken text in quotation marks and identify the speaker. Keep a 5-second line short. For 10 seconds, leave room for pauses and physical action. Describe background sound separately from dialogue.
How to prompt WAN 2.5
A reliable prompt works like a short shot brief. Define the subject, action, scene, camera, visual treatment, and sound in that order. For image-to-video, concentrate on what changes after the first frame.
| Layer | What to describe | Example |
|---|---|---|
| Subject | Person, character, object, or product | A cyclist in a yellow rain jacket |
| Action | One readable movement | Rides through a flooded street |
| Scene | Location, weather, and time | Empty city at blue hour after heavy rain |
| Camera | Framing and one main movement | Low side-tracking shot |
| Lighting | Direction, contrast, and palette | Cool streetlights reflected in water |
| Dialogue | Exact line and delivery | She says quietly, “We keep moving.” |
| Sound | Effects, ambience, or music | Tire spray, distant thunder, low pulse |
| Continuity | What must remain stable | Preserve face, jacket, bicycle, and weather |
Ten seconds is enough for one developed action or two simple beats. Several locations, costume changes, and camera movements in one prompt usually weaken continuity.
WAN 2.5 prompt examples
Rainy cyclist
Low side-tracking shot of a cyclist in a yellow rain jacket riding through a flooded city street at blue hour. Cool lights reflect in the water. Tire spray, distant thunder, and a restrained electronic pulse. Preserve the rider and bicycle throughout.
Street dialogue
Medium close-up of a woman waiting beneath a flickering bus-stop light at night. She looks down the empty road and says quietly, “You said you would come back.” Slow push-in. Light rain and distant traffic. No other voices.
Skincare ad
Macro product film of an amber serum bottle beside sliced citrus and wet leaves. The camera slides slowly across condensation as morning light reaches the label. Sound of water drops, soft glass resonance, airy ambient music. Keep the bottle geometry and label readable.
Coffee spot
Close cinematic sequence in a quiet café. Espresso pours into a white cup, steam curls upward, and a hand places the cup beside an open notebook. Rich machine sound, ceramic contact, soft morning room tone. Warm natural light.
Running shoe
Low tracking shot following a runner under an elevated railway before sunrise. The shoes strike wet pavement in rhythm. Pale blue haze, gritty realism. Synchronized footsteps, breath, and a train passing overhead.
Product presenter
Vertical UGC shot of a creator holding a compact camera in a bright apartment. She smiles and says, “This is the one feature I use every day.” She turns the screen toward the viewer. Natural room tone and one soft notification sound.
Fantasy gate
Wide cinematic shot of a traveler approaching a stone doorway suspended in a glowing forest. The symbols activate as she raises her hand. Slow camera arc, drifting spores, low wind, stone resonance, and distant choral music.
Animated mascot
A small illustrated orange robot jumps out of a delivery box, rolls across the table, and says cheerfully, “Right place, right time.” Preserve the drawing style and proportions. Cardboard rustle, wheel sounds, bright two-note music sting.
Fashion walk
Full-body fashion film in a white concrete gallery. A model in a structured red coat walks toward the camera, stops, and turns in profile. Smooth backward tracking, crisp footsteps, subtle fabric movement, minimal electronic beat.
Food commercial
Overhead view of fresh pasta being finished with parmesan and basil on a dark wooden table. One continuous slow push-in. Detailed grating and plating sounds, quiet restaurant ambience, warm editorial lighting.
Forest wanderer
Behind view of a lone traveler moving through a dark forest as heavy fog curls around the trees. The camera follows in one continuous shot. Saturated fantasy palette, old-film atmosphere, slow suspenseful pacing, branches and distant birds.
Sports interview
Medium shot of a boxer catching his breath beside the ring. He looks at the interviewer and says, “The last round starts in your head.” Slight handheld movement, gym reverb, rope creaks, and distant training impacts.
Music performance
Close performance shot of a singer on a small rooftop at sunset, synchronized to the supplied vocal audio. Natural mouth movement and restrained gestures. Slow clockwise camera arc, wind in clothing, city ambience under the music.
Car reveal
A silver electric car emerges from a dark tunnel into early morning light. Low front tracking shot, then a controlled side reveal. Tire noise, electric motor tone, tunnel reflections, no dialogue. Preserve body shape, wheels, lights, and paint.
Russian dialogue
Крупный план мужчины у окна поезда на рассвете. Он смотрит на отражение и спокойно говорит: «Иногда дорога важнее пункта назначения». Медленный отъезд камеры, стук колёс, тихий гул вагона, без других голосов.
Building a WAN 2.5 workflow in Phygital+
A WAN pipeline can prepare the image, prompt, and sound before generation. Start with a text node for the shot and dialogue. Generate a consistent first frame in an image model when identity or product shape matters. Create or upload the audio, then connect the approved materials to WAN 2.5.
- Write one 5- or 10-second shot with dialogue and sound layers.
- Create or select the starting image for image-to-video.
- Generate or upload voice, music, or effects if custom audio is needed.
- Run several WAN 2.5 branches with one controlled difference.
- Review motion, identity, lip-sync, and sound together.
- Continue the selected output to upscaling or editing.
Keeping every input and branch on the canvas makes it easier to replace a line, camera direction, or first frame without losing the approved setup.
WAN 2.5 pricing in Phygital+
WAN 2.5 uses Phygital credits. The current cost appears in the node before generation and changes with the exposed duration and resolution settings. Test the direction at the lowest practical setting, then move the approved branch to the final output quality.
Common WAN 2.5 problems
The scene feels rushed
Reduce the prompt to one primary action and use 10 seconds when the scene needs dialogue plus movement.
The generated dialogue is unclear
Shorten the line, identify the speaker, describe the delivery, and request no additional speech.
The sound does not match the action
Describe sound at the moment it occurs and keep effects separate from ambience and music.
The character changes
Use image-to-video with an approved first frame and repeat only the essential identity constraints.
The camera becomes chaotic
Specify one main camera movement. Remove stacked commands such as orbit, zoom, pan, and crane in the same short clip.
Uploaded audio gets cut
Match the audio edit to the selected 5- or 10-second duration before generation.
Current limitations
WAN 2.5 is built around short clips. Complex multi-shot stories, several speakers, rapid hand-object interaction, crowded scenes, and long dialogue can reduce consistency. Automatic audio may also invent unwanted speech or music unless the prompt defines the sound clearly.
WAN 2.5 preview is available through hosted Alibaba Cloud endpoints. Do not assume that the open-weight terms of earlier Wan releases apply to this exact model ID.
FAQ
What is WAN 2.5?
WAN 2.5 is Alibaba’s short-form text-to-video and image-to-video model with synchronized native or uploaded audio.
How long are WAN 2.5 videos?
Official preview endpoints support 5- and 10-second output.
What resolutions does WAN 2.5 support?
Alibaba documentation lists 480p, 720p, and 1080p output at 30 fps in MP4 format.
Does WAN 2.5 generate audio?
Yes. It can generate automatic dubbing, effects, music, and ambience, or synchronize video to an uploaded audio file.
Should I use text-to-video or image-to-video?
Use text-to-video for exploration. Use image-to-video when the first frame, character, product, composition, or visual identity is already approved.
Can I upload my own audio?
Yes. Official documentation supports WAV or MP3 audio. Edit it to the selected video duration so important content is not cut.
Can WAN 2.5 create dialogue?
Yes. Put the exact line in quotation marks, name the speaker and delivery, and keep it short enough for the selected duration.
Is WAN 2.5 open source?
Earlier Wan releases have open-weight variants, but WAN 2.5 preview is documented through hosted Alibaba Cloud endpoints. Check the exact model license rather than applying another version’s terms.
How much does WAN 2.5 cost in Phygital+?
The model uses Phygital credits, and the current cost is displayed in the node before generation.
Why use WAN 2.5 in Phygital+?
Phygital+ connects script development, first-frame generation, voice and sound preparation, WAN video branches, editing, and upscaling on one canvas.