When I looked through ByteDance’s first Seedance 2.5 examples, the extra duration immediately stood out. Thirty seconds gives the model enough time to establish a location, introduce a subject, move through several actions, and reach an ending.
That changes how the prompt should be written. A description of one shot is rarely enough. Seedance 2.5 responds better when the prompt explains how the scene develops over time.
The model also accepts a large collection of reference material. According to the official announcement, one generation can use up to 30 images, 10 video clips, and 10 audio clips. References can define characters, objects, visual language, camera movement, pacing, and sound.
In this guide, we will cover version differences, text-to-video, image-to-video, reference workflows, timecoded prompting, editing, consistency, and production pipelines in Phygital+.
Seedance 2.5 compared with Seedance 2.0
Seedance 2.0 established ByteDance’s multimodal approach to video generation. It supports text, images, audio, and video as inputs and generates visuals and sound together.
Seedance 2.5 expands this approach around longer sequences, larger reference sets, and more precise editing. ByteDance officially released it on July 31, 2026.
| Version | Best use | Why I would use it |
|---|---|---|
| Seedance 2.0 | Short video concepts and quick tests | Established multimodal generation with clips up to 15 seconds |
| Seedance 2.5 | Longer scenes and developed narratives | Generates up to 30 seconds and maintains more context across the sequence |
| Seedance 2.5 with references | Character, product, and campaign consistency | Images, video, and audio can define multiple parts of the result |
| Seedance 2.5 editing | Improving an existing generation | Supports targeted changes to selected moments and camera decisions |
Seedance 2.5 supports durations from 4 to 30 seconds. Current BytePlus documentation lists 480p and 720p output for this model. Higher resolutions available in other Seedance configurations should not be assumed to apply to 2.5.
What Seedance 2.5 is good at
Longer scenes
Seedance 2.5 can generate up to 30 seconds of synchronized video and audio in one pass. The model can use that time for connected actions, camera transitions, and changes in location.
It also supports multi-round extension. An approved clip can become the starting point for the next section while preserving its main characters, environment, sound, and pacing.
Multimodal references
Images can define a character, product, location, wardrobe, palette, or composition. Video references can communicate movement and camera language. Audio references can establish music, voice, rhythm, or environmental sound.
- Up to 30 reference images
- Up to 10 reference videos
- Up to 10 reference audio clips
These are model-level limits. The exact number exposed in a particular interface or API workflow can vary.
Timeline control
Timecodes help organize longer generations. You can describe what should happen during 0–5s, 6–12s, and 13–20s, including action, camera movement, scene transitions, and sound.
Editing selected moments
Seedance 2.5 supports timestamp-based editing of generated or referenced video. Typical requests include changing an action, adjusting camera movement, replacing a background, modifying pacing, or adding an object. ByteDance also documents green-screen, camera-perspective, and reference-based editing workflows.
Character and object consistency
Reference images give the model a stable visual description of the subject. For a character, this can include the face, hairstyle, outfit, body proportions, and accessories. For a product, references can establish its shape, materials, label, colors, and scale.
How to prompt Seedance 2.5
I would approach Seedance 2.5 as a sequence planner. Begin with the format and duration, establish the subject and location, divide events into time ranges, describe camera and sound inside each range, and finish with continuity requirements.
| Component | What to describe | Example |
|---|---|---|
| Format | Duration, aspect ratio, visual treatment | 20-second cinematic product film in 16:9 |
| Subject | Main character, object, or product | A dark green glass perfume bottle |
| Scene | Location and environment | Wet volcanic stone surrounded by mist |
| Timeline | Order and timing of actions | 0–5s close-up, 6–12s reveal, 13–20s orbit |
| Camera | Framing and movement | Macro slide, slow pullback, clockwise arc |
| Lighting | Direction, color, and intensity | Narrow warm light crossing the bottle |
| Audio | Dialogue, ambience, music, effects | Soft wind, glass resonance, low ambient score |
| Continuity | Details that must remain stable | Preserve bottle shape, label, cap, and color |
Writing prompts with timecodes
Timecodes should describe meaningful sections of the scene. A useful 20-second structure is 0–5s for the setup, 6–12s for development, 13–17s for a reveal, and 18–20s for the final frame.
Keep each block readable. Assign one main action and one primary camera movement to each time range.
Seedance 2.5 prompt examples
Cinematic: Desert courier
A 24-second cinematic sequence in 16:9. A motorcycle courier crosses a pale desert at sunrise carrying a sealed metal case. 0–7s: wide tracking shot alongside the motorcycle. 8–16s: the rider approaches an abandoned communications tower as the camera rises. 17–24s: the rider stops, removes the case, and looks toward a distant storm. Low wind, engine vibration, restrained electronic score.
Cinematic: Night train
A 20-second atmospheric thriller inside an almost empty night train. 0–6s: follow a woman walking through the carriage while passing lights sweep across the windows. 7–14s: she notices a handwritten envelope and reaches for it. 15–20s: the train enters a tunnel, the lights flicker, and the envelope disappears. Realistic train ambience and quiet suspense.
Cinematic: Coastal rescue
A 25-second realistic rescue scene during a storm. 0–8s: a coastguard runs across a wet pier toward an emergency boat. 9–18s: the boat cuts through rough waves in a low side-tracking shot. 19–25s: a flare appears in the distance. Preserve the same person, uniform, boat, weather, and lighting.
Cinematic: Museum after hours
A 20-second fantasy sequence inside a closed museum. A marble bird sculpture gradually comes to life. Begin with a wide shot, move into a close-up as stone feathers shift, then follow the bird through the gallery. Marble dust falls naturally. Quiet room ambience and soft orchestral music.
Product ad: Skincare serum
A 15-second premium skincare film in 9:16. A translucent serum bottle stands beside sliced yuzu and wet green leaves. Start with a macro shot of condensation, pull back as sunlight moves across the bottle, then finish with a centered product frame. Preserve the exact bottle shape, cap, label, and liquid color from the reference.
Product ad: Running shoe
A 20-second sports commercial in 16:9. 0–5s: extreme close-up of a runner tightening the laces. 6–13s: low tracking shot following the shoe across wet pavement. 14–20s: the runner accelerates under an elevated railway as the camera pulls into a wide city view. Crisp footsteps and distant train sound.
Product ad: Coffee machine
A 16-second editorial product demo. Show a compact chrome coffee machine in a warm modern kitchen. Move from a clean product shot to a close-up of beans grinding, espresso pouring, and steam rising. Keep the machine geometry and control layout unchanged.
Product ad: Electric car
A 24-second automotive film featuring the same silver electric car from the supplied references. Show the car leaving an underground parking structure, moving through a rainy city, and arriving on a rooftop at dawn. Preserve the body shape, wheels, headlights, paint, and interior.
Social: Fashion transition
A 12-second vertical fashion clip. A woman walks toward the camera through a white studio while her outfit changes at three moments: casual denim at 0–4s, structured black tailoring at 5–8s, metallic eveningwear at 9–12s. Keep her face, hairstyle, movement, and framing consistent.
Social: Restaurant menu
A 15-second overhead food video in 9:16. Ingredients enter the frame one by one and assemble into a finished pasta dish. Add realistic cooking sounds and keep the tabletop, plate, utensils, and lighting fixed.
Social: Workspace transformation
A 15-second social clip showing a messy home desk becoming an organized creative workspace. The camera performs one slow push-in while objects rearrange in a clear sequence. Preserve the room architecture and daylight direction.
Performance: Rooftop singer
A 25-second live performance on a city rooftop at sunset. Use the supplied character reference for the singer and the audio reference for timing. Begin with a close portrait, circle into a medium shot during the first phrase, then reveal the skyline and band. Preserve identity, clothing, microphone, and vocal timing.
Performance: Contemporary dance
A 20-second contemporary dance film inside a large concrete hall. A single dancer crosses through narrow shafts of light. Move from a wide architectural view to a side-tracking medium shot and finish overhead. Use the reference video for movement rhythm.
Performance: Studio drummer
A 15-second energetic music clip. A drummer performs under red and blue studio lighting while the camera moves from the cymbal to the performer’s face and then pulls out to reveal the full kit. Synchronize visible strikes with the supplied audio.
Consistency: Delivery robot
A 20-second animated sequence featuring the same small orange delivery robot in every shot. The robot leaves a warehouse, crosses a busy square, avoids a puddle, and delivers a package to a florist. Preserve its proportions, screen face, wheels, orange panels, and package.
Consistency: Explorer story
A 30-second fantasy sequence using supplied character and environment references. 0–8s: the explorer enters a glowing forest. 9–18s: she finds a suspended stone doorway. 19–26s: it opens onto a snowy mountain. 27–30s: hold on her face as cold air moves through her hair. Preserve her identity, coat, backpack, and pendant.
Consistency: Brand mascot
A 15-second commercial featuring the supplied illustrated mascot. The mascot jumps out of a cereal box, runs across the breakfast table, and points toward the product label. Maintain the original illustration style, colors, face, and proportions.
Editing: Camera edit
Edit the supplied video from 0–15 seconds. Keep the people, actions, environment, lighting, and visual style unchanged. Adjust only the camera: slow push-in from 0–5s, smooth lateral track from 6–10s, and a gentle arc around the subject from 11–15s.
Editing: Background replacement
Use the supplied studio performance video. Keep the performer, choreography, timing, clothing, and camera movement unchanged. Replace the green-screen background with a large industrial hall filled with soft morning haze. Match the new lighting and shadows to the performer.
Editing: Product environment
Edit the supplied product video. Preserve the product, label, materials, reflections, and original movement. Replace the white studio with a dark stone environment, add a narrow warm light from the left, and introduce subtle mist behind the product.
Using references effectively
References should have clear responsibilities. One image can define the face, another the outfit, a third the location, while a video establishes camera movement and an audio clip establishes rhythm.
Naming these responsibilities inside the prompt helps the model understand how each asset should influence the result. Remove duplicated or contradictory references when identity begins to drift.
Building a Seedance 2.5 workflow in Phygital+
A useful Seedance pipeline begins before the video node. Start with a text model to develop the idea and scene order, create the character, product, or first frame in an image model, then feed approved references into Seedance 2.5 and generate several motion directions.
- Write the concept and timecoded shot sequence with a text model.
- Generate character, product, and location references.
- Connect the approved references to Seedance 2.5.
- Create several branches with different camera or pacing decisions.
- Select the strongest result.
- Extend or edit the chosen clip.
- Send the final video into an upscaler or additional audio step.
This keeps the approved visual foundation visible throughout the project. New camera and pacing options can branch from the same source material.
Seedance 2.5 pricing in Phygital+
Seedance 2.5 uses Phygital credits. The exact cost appears inside the node before generation and depends on the selected duration and available settings.
Develop the references and timecoded structure before starting the longest generation. Short exploratory branches can validate the character, camera, and visual direction first.
Common Seedance 2.5 problems
The middle of the video loses direction
Divide the sequence into three or four time ranges and assign one main action to each.
The character changes between shots
Use a clean identity reference and repeat the important continuity requirements at the end of the prompt.
Camera instructions are ignored
Reduce the number of simultaneous commands and assign one primary camera movement to each time block.
References produce conflicting results
Give every reference a clear purpose and remove images that disagree about the same detail.
The 30-second clip feels slow
Define a setup, development, transition, and final beat so the duration has enough events to support it.
Editing changes too much
State the exact time range and list every element that must stay unchanged.
Current limitations
ByteDance acknowledges that complex physics and interactions between multiple subjects still have room for improvement. Crowded scenes, fast contact, intricate hand movement, and several simultaneous actions can introduce visual errors.
Keep the central action readable, use references for important details, and introduce complexity gradually.
FAQ
How long can a Seedance 2.5 video be?
Seedance 2.5 supports clips from 4 to 30 seconds in a single generation. It also supports multi-round extension for longer sequences.
What is the difference between Seedance 2.0 and Seedance 2.5?
Seedance 2.0 supports shorter generations up to 15 seconds. Seedance 2.5 extends this to 30 seconds and adds broader multimodal reference capacity, stronger long-form continuity, timestamp control, and improved editing.
Does Seedance 2.5 generate audio?
Yes. Seedance 2.5 uses joint audio-video generation, allowing sound and visuals to be created together.
How many references can Seedance 2.5 use?
The official model announcement lists up to 30 images, 10 video clips, and 10 audio clips. Platform and API limits can vary by workflow.
What resolution does Seedance 2.5 support?
Current BytePlus documentation lists 480p and 720p for Seedance 2.5. Other Seedance models may expose higher resolutions.
Should I use text-to-video or image-to-video?
Use text-to-video for visual exploration. Use image-to-video when the character, product, composition, or starting frame is already approved.
How do timecodes work in Seedance prompts?
Timecodes divide the prompt into chronological sections. Each section can define its own action, framing, camera movement, and sound.
Can Seedance 2.5 edit an existing video?
Yes. It supports targeted editing, camera perspective changes, green-screen workflows, and reference-based modifications.
How do I keep the same character throughout the video?
Use a clear character reference, avoid conflicting images, keep wardrobe descriptions stable, and list the identity details that must remain unchanged.
Why use Seedance 2.5 in Phygital+?
Phygital+ connects script development, reference-image generation, Seedance video, branching, editing, and upscaling on one canvas.