Seedance 2.0 Collection

Seedance 2.0 × VidModel

Seedance 2.0 is ByteDance's cinematic video generation engine — native audio sync, up to 4K resolution, and four generation modes covering image, text, and multi-reference inputs. Coming soon via VidModel's unified API.

Model Group

Cinematic video generation across four input modes, with native audio and 4K output

Seedance 2.0 is built for production-quality video — accurate motion, consistent subject identity across frames, and native audio generation baked in. Choose from image-to-video, text-to-video, standard reference-to-video, or the fast reference-to-video tier for high-throughput pipelines.

4

generation modes

4K

max output resolution

Native

audio generation

Included models

Four models, one integration — pick by input type and throughput target.

Seedance 2.0 Image to Video

Animate still images into cinematic video with physics-accurate motion, native audio sync, and up to 4K resolution — powered by ByteDance.

IMAGE INPUTStandard
Try it

Seedance 2.0 Text to Video

Generate cinematic video from text prompts with automatic camera planning, multimodal reference support, and native audio sync — up to 4K.

TEXT INPUTStandard
Try it

Seedance 2.0 Reference to Video

Multi-reference video generation with strong character and style consistency — supply images, video clips, and audio to anchor identity across shots.

REFERENCE INPUTStandard
Try it

Seedance 2.0 Fast Reference to Video

Speed-optimized Seedance R2V for rapid iteration — same multimodal reference inputs at significantly faster generation time and lower cost.

REFERENCE INPUTStandard
Try it

Pick the right model

Four generation modes, one API pattern — pick by input type and throughput target.

ModelInputSpeedBest For
Source imageStandardCinematic portrait animation, branded video
Text promptStandardPrompt-driven scenes, concept generation
Reference imagesStandardMulti-reference composition, ensemble scenes
Reference imagesFastHigh-volume pipelines, personalization at scale
PRODUCTION PLAYBOOKS

Match each Seedance 2.0 model to a real video workflow

Image, text, reference, or fast reference — one API pattern handles all four generation modes.

01
Image to Video

Cinematic portrait animation

Animate a source image into a cinematic clip with accurate subject motion, consistent identity across frames, and native audio layered in. Seedance I2V handles complex scene lighting and background preservation — the output character moves expressively rather than rigidly, suitable for production-quality branded video.

02
Text to Video

Prompt-driven scene generation

Describe any scene, character, or action in a prompt and generate a full cinematic video with no source image required. Seedance T2V handles complex motion descriptions, multi-subject scenes, and specific lighting or atmosphere briefs — a direct path from creative brief to finished video clip.

03
Reference to Video

Multi-reference scene composition

Supply multiple reference images to anchor subject appearance, style, or scene composition — then generate video that holds all reference details consistently across frames. Seedance R2V is suited for ensemble scenes, brand-consistent character video, or any workflow where subject fidelity across a clip is the primary constraint.

04
Fast Reference

High-throughput reference pipelines

The Fast Reference tier cuts generation latency significantly compared to standard R2V while maintaining the same multi-reference subject fidelity. Kick off batches concurrently with Promise.all — the only real ceiling is your account rate cap, not sequential generation time. The right tier for personalization pipelines and high-volume content operations.

Seedance2.0 API
1
Image to Video
seedance2.0-i2v
2
Text to Video
seedance2.0-t2v
3
Reference to Video
seedance2.0-r2v
4
Fast Reference
seedance2.0-r2v-fast

Native audio generation

Seedance 2.0 generates ambient sound or voiceover-ready audio alongside video in the same request — no separate audio pipeline or post-processing required.

Up to 4K output

Generate broadcast-ready video at up to 4K resolution — suitable for premium digital delivery, OTT, or high-fidelity brand content.

Four modes, one integration

Image, text, reference, and fast reference all share the same task endpoint and polling pattern — switch generation modes without changing your integration.