
Seedance 2.0 × VidModel
Seedance 2.0 is ByteDance's cinematic video generation engine — native audio sync, up to 4K resolution, and four generation modes covering image, text, and multi-reference inputs. Coming soon via VidModel's unified API.
Cinematic video generation across four input modes, with native audio and 4K output
Seedance 2.0 is built for production-quality video — accurate motion, consistent subject identity across frames, and native audio generation baked in. Choose from image-to-video, text-to-video, standard reference-to-video, or the fast reference-to-video tier for high-throughput pipelines.
4
generation modes
4K
max output resolution
Native
audio generation
Included models
Four models, one integration — pick by input type and throughput target.

Seedance 2.0 Image to Video
Animate still images into cinematic video with physics-accurate motion, native audio sync, and up to 4K resolution — powered by ByteDance.

Seedance 2.0 Text to Video
Generate cinematic video from text prompts with automatic camera planning, multimodal reference support, and native audio sync — up to 4K.

Seedance 2.0 Reference to Video
Multi-reference video generation with strong character and style consistency — supply images, video clips, and audio to anchor identity across shots.

Seedance 2.0 Fast Reference to Video
Speed-optimized Seedance R2V for rapid iteration — same multimodal reference inputs at significantly faster generation time and lower cost.
Pick the right model
Four generation modes, one API pattern — pick by input type and throughput target.
| Model | Input | Speed | Best For |
|---|---|---|---|
| Source image | Standard | Cinematic portrait animation, branded video | |
| Text prompt | Standard | Prompt-driven scenes, concept generation | |
| Reference images | Standard | Multi-reference composition, ensemble scenes | |
| Reference images | Fast | High-volume pipelines, personalization at scale |
Match each Seedance 2.0 model to a real video workflow
Image, text, reference, or fast reference — one API pattern handles all four generation modes.
Cinematic portrait animation
Animate a source image into a cinematic clip with accurate subject motion, consistent identity across frames, and native audio layered in. Seedance I2V handles complex scene lighting and background preservation — the output character moves expressively rather than rigidly, suitable for production-quality branded video.
Prompt-driven scene generation
Describe any scene, character, or action in a prompt and generate a full cinematic video with no source image required. Seedance T2V handles complex motion descriptions, multi-subject scenes, and specific lighting or atmosphere briefs — a direct path from creative brief to finished video clip.
Multi-reference scene composition
Supply multiple reference images to anchor subject appearance, style, or scene composition — then generate video that holds all reference details consistently across frames. Seedance R2V is suited for ensemble scenes, brand-consistent character video, or any workflow where subject fidelity across a clip is the primary constraint.
High-throughput reference pipelines
The Fast Reference tier cuts generation latency significantly compared to standard R2V while maintaining the same multi-reference subject fidelity. Kick off batches concurrently with Promise.all — the only real ceiling is your account rate cap, not sequential generation time. The right tier for personalization pipelines and high-volume content operations.
Native audio generation
Seedance 2.0 generates ambient sound or voiceover-ready audio alongside video in the same request — no separate audio pipeline or post-processing required.
Up to 4K output
Generate broadcast-ready video at up to 4K resolution — suitable for premium digital delivery, OTT, or high-fidelity brand content.
Four modes, one integration
Image, text, reference, and fast reference all share the same task endpoint and polling pattern — switch generation modes without changing your integration.