Doubao Seedance 2.5 AI Video Generator - 30-Second One-Take Storytelling & Multimodal Reference
Doubao Seedance 2.5 (Dreamina Seedance 2.5) is the new-generation video creation model released by ByteDance Seed on July 31, 2026. Building on the unified multimodal audio-video joint-generation architecture of Seedance 2.0, it breaks through in three areas: 30-second single-pass generation with multi-round extension, fully upgraded multimodal reference with up to 30 images, 10 videos, and 10 audio clips, and precise, stable editing with timestamp-level control. Seedance 2.5 moves AI video from "generating a clip" to "completing a creative work".
Use cases: Film & Short Drama | Advertising & Product Marketing | Music & Stage Performance | Education & Explainers | Industrial Simulation
Seedance 2.5 Core Features
The short creative film below was produced end-to-end by Seedance 2.5:
🎬 30-Second Long-Form Storytelling
Seedance 2.5 extends single-pass video generation from 15 to 30 seconds and strengthens storytelling across the whole clip:
- Complete Story Arcs: Within 30 seconds, the model organizes multiple logically connected shots so the story unfolds through setup, development, turning point, and resolution — not just an extension of a single moment
- Smooth One-Take Camera Work: More natural camera transitions, with the main subject staying stable across cuts and audio and visuals staying in sync
- Real-World Physics: Motion such as flowing sleeves or running crowds follows more natural arcs and trajectories
- Less Artificial Look: Textures, skin, eyes, lighting, and color saturation are systematically optimized, with fewer uncontrolled subtitles and background music, for results closer to live-action footage
T2V prompt: One-take handheld gimbal tracking shot. The camera slowly pushes in through a gap in a heavy red curtain and enters a warm-toned backstage dressing room. A young female singer, with her back to the camera, is adjusting her earpiece as a staff member reminds her it's time to go on. She turns toward the camera and starts singing citypop. The camera pulls back and tracks her as she passes through the curtain into a dim backstage corridor, interacting naturally with her dancers along the way; one staff member hands her a microphone. She and the dancers then step onto the stage, and the camera arcs around to the back, gradually revealing the red-and-black stage design, LED screens, spotlights, haze, and reflective floor. The camera finally pulls out to a wide shot of the arena, showing the packed audience, light boards, glow sticks, and cheering crowd.
R2V prompt: 16:9 widescreen, cinematic texture, single continuous take, smooth camera movement, no cuts. Scene reference: @Image 4. 0–5s: Open with a close-up of the Overlord from @Image 2. The camera slowly circles his upper body and transitions into a medium shot. The Overlord spins and turns, his body and back flags sweeping quickly past the lens to form a natural occlusion, and the camera follows through to Consort Yu's side in @Image 1. 6–10s: The camera steadily circles Consort Yu in a medium shot, following her water sleeves through the arc. She raises her arm, flicks her wrist, unfurls the sleeves, and half-turns, then draws the sleeves back, holds the pose, and looks sideways toward the Overlord. 11–20s: The male warrior from @Image 3 enters with an aerial flip. The Overlord takes center stage while the warrior advances and retreats on the opposite side. The camera slowly pulls back to a full stage view, and all three face the audience and strike a synchronized Peking opera finale pose.
🔗 Multi-Round Video Extension
Append subsequent shots to an existing Seedance 2.5 output and keep the story going:
- Consistent Characters & Scenes: Main characters, environments, visual style, and sound effects stay consistent across every extension
- Continuous Narrative Pacing: New shots connect naturally to the previous clip instead of restarting the scene
- Multi-Minute Content: Build videos lasting several minutes with a unified audiovisual language, reducing the effort of splitting clips, splicing footage, and fixing transitions
R2V prompt: Extend the video. Continue from the visuals and subjects in @Video 1 and generate another 30-second clip, keeping the character subjects, scene, visual style, and sound effects consistent. The little boy runs along the train carriage holding a soccer ball. When the subway stops, the side door opens and he immediately dashes out, with the male lead chasing after him. The two run across the platform and out onto the street, startling passersby along the way. The male lead finally catches up and grabs him. The boy looks up, aggrieved. The male lead's anger slowly fades; he pats the boy's head and shows a helpless smile.
📎 Up to 50 Multimodal References
Seedance 2.5 accepts up to 30 images, 10 video clips, and 10 audio clips in a single generation:
- Comprehensive Understanding: Composition, scenes, styles, characters, and props from every reference are applied as instructed
- Multi-Character Consistency: In group scenes such as concerts or ensemble stories, the appearance and voice of each character stay stable
- Motion & Creative Reference: Better understanding of the intent and camera language behind reference videos, upgrading motion transfer to creative interpretation
R2V prompt: A 30-second concert sequence in 16:9 landscape, with cinematic realism, authentic concert hall lighting, warm golden stage lighting, and the atmosphere of a formal classical concert. Use @Image 1 for the venue. Reference @Image 2 for the pianist, @Image 3 for the cello, @Image 4 for the violin, and @Image 5 for the lead vocalist. Reference @Images 6 to 10 for the rest of the orchestra, @Images 11 to 14 for the choir, and @Images 15 to 18 for the audience. Open with a high-angle wide shot of the full concert hall. The pianist strikes the keys, and the lead vocalist steps into the spotlight and begins singing. The camera moves across the violin, cello, and orchestra; in the latter part, the choir joins in. In the closing shot, the camera pulls back, the singing ends, and the audience joins in the applause.
🧊 Clay Render Reference & Lighting Control
Block out your shot with textureless 3D models and let Seedance 2.5 render the final look:
- Spatial Blocking: Define spatial structure, character poses, motion paths, and camera positions with a clay render
- Predictable Composition: Complex shots follow the composition and blocking you planned
- Physically Based Lighting: Spatial information from the clay render drives realistic light direction, color temperature, intensity, and shadows
R2V prompt: Refer to @Clay Render 1 for camera movement, pacing, shot-size transitions, subject trajectory, and blocking. Refer to @Image 2 for character design, scene, materials, lighting, color, and fairy-tale atmosphere, and render the clay render as a dreamy, warm 3D animated short with a childlike fantasy feel. The story unfolds as follows: flight through a fantasy sky → mythical beasts flying alongside through a sea of clouds → a dive into the ocean → weaving through the deep with manta rays → passing through a mirrored rift in spacetime → picking stars from the cosmos → transforming back into the bedroom → father tucking in the blanket → the picture book closes and holds on the final frame.
R2V prompt: Reference the camera work, composition, shot scale, spatial relationships, part positions, model structure, assembly order, and motion paths from @Clay Render 1. Reference the materials, lighting, color, reflections, and atmosphere from @Image 1, and turn the clay render into a high-end, photorealistic car assembly sequence.
✂️ Precise, Timestamp-Level Editing
Refine details without regenerating from scratch:
- Timestamp Control: Specify what happens in each time range — story beats, camera perspective, movement, and rhythm
- Targeted Edits: Modify characters, actions, sound, or plot within a specific segment while keeping everything before and after consistent
- Green Screen Editing: Replace backgrounds and tell entirely different stories while keeping the subject intact, with clothing, hair, gait, and lighting adapted to the new environment
- Camera Editing: Keep characters, actions, and visual style unchanged and redesign only the camera movement
R2V prompt: Using @Video 1, render the green-screen background, obstacles, wardrobe, and supporting characters. 0–4s: outdoor training, replace the obstacles with rocks, bricks, tires, and wooden crates. 4–10s: locker room, friends offering encouragement. 10–15s: international match, replace the training poles with original defenders and a goalkeeper, and the protagonist scores. Overall photorealistic, cinematic quality.
R2V prompt: Edit @Video 1. Keep the characters, actions, and visual style unchanged. Adjust only the camera movement. 0–4s, a micro-FPV move skims tightly past the pan, then follows the popping toast and whip-pans to the coffee; 4–7s, push in and track laterally along the rim of the pan, following the fried egg as it flips up and lands back in place; 7–11s, rapidly rise to a top-down view, then descend at a steady pace, sweeping across the plate and keys; 11–15s, use a handheld close-up to follow the hands with a fast lateral whip, then push in on the breakfast and pull back to a medium two-shot.
Seedance 2.5 Specifications
| Item | Seedance 2.5 |
|---|---|
| Developer | ByteDance Seed |
| Release Date | July 31, 2026 |
| Architecture | Unified multimodal audio-video joint generation |
| Duration | 4–30 seconds per generation, multi-round extension |
| Resolution | 480p / 720p / 1080p |
| Image References | Up to 30 |
| Video References | Up to 10 |
| Audio References | Up to 10 |
| Output | Video with synchronized audio |
| Advanced Control | Clay render reference, timestamp editing, green screen editing, camera editing |
Seedance 2.5 vs Seedance 2.0
| Feature | Seedance 2.5 | Seedance 2.0 |
|---|---|---|
| Max Duration per Generation | 30 seconds | 15 seconds |
| Resolution | Up to 1080p | Up to 720p |
| Video Extension | Multi-round, consistent characters and pacing | Forward extension, prequel, track completion |
| Reference Materials | 30 images + 10 videos + 10 audio clips | Image, video, and audio references |
| Clay Render Reference | Supported, with lighting control | — |
| Editing | Timestamp-level editing, green screen, camera, reference editing | Subject replacement, object add/remove, inpainting |
| Visual Quality | Optimized textures, skin, and lighting for a less artificial look | Precise instructions and realistic materials |
| Best For | Long-form stories, music videos, multi-character scenes, professional production | Short ads, product demos, precise single-shot control |
Applicable Scenarios & Use Cases
🎬 Film, Short Drama & Music Videos
- One-Take Sequences: Generate 30-second continuous shots with complete story arcs for short dramas and trailers
- Stage & Music Performances: Keep singers, dancers, orchestras, and choirs consistent in multi-character concert scenes
- Long-Form Stories: Chain multiple 30-second extensions into multi-minute narratives
📱 Advertising & Product Marketing
- Product Replacement: Swap products, styles, faces, lighting, or backgrounds while keeping original motion and composition
- Brand Consistency: Use dozens of reference images to lock product details, characters, and brand assets
- Green Screen Production: Shoot once on green screen and render multiple scenes and stories for campaign variations
🎓 Education & Explainers
Seedance 2.5 turns lessons, historical context, and scientific principles into vivid, immersive video — lowering the barrier for teachers to produce instructional content.
R2V prompt: Expressive Eastern painterly style. A street scene in Lin'an during the Southern Song dynasty. Several children run and shout through the bustling street, chanting, "I turn around, and there he is, where the lantern lights grow dim." The camera follows the children as they run, sweeping past the lively street. The camera then tilts up to reveal Xin Qiji from @Image 1. Xin Qiji turns his head, and in the distance stands a man among the fading lantern lights. The shot stays continuous throughout.
🏭 Industrial & Professional Production
- Industrial Simulation: Turn clay renders into photorealistic assembly sequences, process training, and equipment demos
- Synthetic Training Data: Generate high-quality video data for robot perception and manipulation training
- Autonomous Driving: Simulate long-tail scenarios such as extreme weather and complex traffic for testing and training
How to Create Videos with Seedance 2.5
1. Access the Platform
Visit the AIGCVA App Center and select Seedance 2.5.
2. Prepare Your Inputs
- Text prompt: Describe the story, shots, camera movement, and sound — use timestamps for precise pacing
- Reference images: Characters, scenes, products, and style (up to 30)
- Reference videos: Motion, camera language, clay renders, or footage to edit (up to 10)
- Reference audio: Voice, music, or sound effects (up to 10)
3. Choose Video Settings
- Duration: 4–30 seconds per generation
- Resolution: 480p, 720p, or 1080p
- Mode: Text-to-video or reference-to-video
4. Generate, Extend & Edit
- Click Generate to create your first 30-second clip
- Extend the result to continue the story with consistent characters and scenes
- Edit specific time ranges, the camera, or the green screen background without starting over
Seedance 2.5 Best Practices
- Write in timestamps: Break the prompt into segments such as "0–5s", "6–10s", and "11–20s" to control what happens when
- Label every reference: Refer to inputs explicitly as @Image 1, @Video 1, or @Clay Render 1, and state which element each reference controls
- Plan the story arc: Use the full 30 seconds for setup, development, and resolution instead of a single action
- Block complex shots with clay renders: For precise camera paths or multi-subject blocking, a clay render gives the most predictable result
- Edit instead of regenerating: When one segment is off, use timestamp editing to fix only that part
Seedance 2.5 Prompt Examples
30-Second Product Story
30-second one-take commercial, 16:9, cinematic realism, soft morning light.
0–8s: The camera pushes in through a kitchen window; a young woman grinds coffee beans,
the sound of the grinder and birdsong in the background.
9–18s: The camera follows her to the balcony as she pours the coffee from @Image 1 into a glass cup,
steam rising, the brand logo on the bag clearly visible.
19–30s: She takes a sip and smiles; the camera pulls back to reveal the city skyline at sunrise,
and a warm female voice says: "Start your day slowly."Multi-Reference Character Scene
Use @Image 1 for the cafe interior. The barista is @Image 2, the customer is @Image 3,
and the cat on the counter is @Image 4. Background music follows @Audio 1.
The customer walks in and greets the barista, the cat stretches and jumps down,
the barista hands over a latte with heart-shaped latte art. Keep all three characters consistent.Video Extension
Extend @Video 1 by 30 seconds, keeping the characters, scene, visual style, and sound consistent.
The rain stops, the girl closes her umbrella and walks out of the alley into a busy night market;
the camera tracks her from behind, then circles to a front medium shot as she stops at a lantern stall.Timestamp Editing
Edit @Video 1. Keep the characters, camera movement, and visual style unchanged.
Only change 6–10s: replace the red sports car with a white vintage convertible,
and change the background sound from traffic noise to seaside waves.Seedance 2.5 FAQ
What is Seedance 2.5?
Seedance 2.5 is the new-generation video creation model released by ByteDance Seed on July 31, 2026. It builds on the unified multimodal audio-video joint-generation architecture of Seedance 2.0 and delivers major breakthroughs in long-form storytelling, multimodal reference, and editing — moving from generating a clip to completing a creative work.
How long can a Seedance 2.5 video be?
Seedance 2.5 generates up to 30 seconds of high-quality audio-video in a single pass, doubling the 15-second limit of Seedance 2.0, with selectable durations from 4 to 30 seconds. It also supports multi-round extension that keeps characters, scenes, and pacing consistent, so you can build multi-minute content with a unified audiovisual language.
How many reference materials does Seedance 2.5 support?
A single Seedance 2.5 generation accepts up to 30 images, 10 video clips, and 10 audio clips — 50 multimodal references in total. The model understands composition, scenes, style, characters, and props across all materials and can preserve the appearance and voice of multiple characters in group scenes.
What is clay render reference in Seedance 2.5?
Clay render reference lets you block out spatial structure, character poses, motion paths, and camera positions with textureless 3D models. Seedance 2.5 follows that structure to generate the video, and uses the spatial information to render physically plausible lighting, including light direction, color temperature, intensity, and shadows.
What editing capabilities does Seedance 2.5 offer?
Seedance 2.5 supports timestamp-level control for targeted editing of characters, actions, sound, or plot in specific segments while keeping the rest consistent. It also improves green screen editing, camera perspective and movement editing, and reference-based editing for film and advertising workflows.
How is Seedance 2.5 different from Seedance 2.0?
Seedance 2.5 extends single-pass generation from 15 to 30 seconds, adds multi-round extension, expands references to 30 images, 10 videos, and 10 audio clips, and introduces clay render reference, lighting control, timestamp-level editing, and stronger green screen editing. It also reduces the overly artificial look often seen in AI-generated video.
Is Seedance 2.5 free to use on AIGCVA?
Seedance 2.5 provides free usage quota on the AIGCVA platform. Registered users receive free generation credits, and high-frequency users can purchase VA Coins for additional quota.
Can videos generated with Seedance 2.5 be used commercially?
Yes, videos generated through the AIGCVA platform support commercial use. Please refer to the platform's terms of service for specific commercial terms.
Seedance 2.5 Technical Advantages Summary
- 30-Second One-Take Generation: Complete story arcs with multiple connected shots in a single pass
- Multi-Round Extension: Multi-minute content with consistent characters, scenes, and pacing
- 50 Multimodal References: Up to 30 images, 10 videos, and 10 audio clips per generation
- Clay Render Control: Plan blocking and camera paths in 3D with physically based lighting
- Timestamp-Level Editing: Precise edits to specific segments without regenerating
- Professional Editing Suite: Green screen, camera perspective, and reference-based editing
- Cinematic Realism: Optimized textures, skin, and lighting for a less artificial look
Start Creating with Seedance 2.5
Tell a complete story in one take with ByteDance Seed's latest video model:
🎬 Try Seedance 2.5 for Free Now