
VidGen × VidModel
VidGen is built around human subjects — delivering exceptional body structure and portrait accuracy across complex action scenes, with background audio support. Access T2V, I2V, and Video Extend through VidModel's unified API.
Exceptional human body and portrait rendering, across three generation modes
VidGen is optimized for human subjects — producing accurate body structure, natural movement, and facial detail even in complex action scenes. Choose from text-to-video, image-to-video, or video extension, all on the same API pattern.
3
included models
20s
max video length
Text + Image
input modes
Included models
Pick the model by input type and output goal. Every card maps to a production API workflow.

VidGen 1.0 Image to Video
Animate images into video with accurate human body structure and natural movement — strong in complex action scenes with background audio support.

VidGen 1.0 Video Extend
Extend VidGen-generated video up to 20 seconds while preserving body structure, character consistency, and scene continuity.

VidGen 1.0 Text to Video
Generate video from text prompts with exceptional human body structure and portrait accuracy — ideal for complex action scenes and audio-rich content.
Pick the right model
All three models share the same request pattern — pick by input type and generation goal.
| Model | Input | Max Output | Best For |
|---|---|---|---|
| Portrait image | 20s | Spokesperson videos, portrait animation | |
| Text prompt | 20s | Ad campaigns, scripted character scenes | |
VidGen 1.0 Video ExtendUtility | Existing video | +20s | Scene extension, continuations |
Put VidGen's human-optimized output to work
Three models, one API pattern — pick by input type and generation goal.
Spokesperson videos
Animate a portrait into a talking, moving character for product pitches, feature walkthroughs, or brand introductions. VidGen's human-body specialization means the output holds accurate posture, limb placement, and natural facial expression through the full clip — without the drift typical of general-purpose video models.
Human-led ad campaigns
Generate people-centered video concepts from text prompts at speed — rapid variant creation without actors, scheduling, or a production shoot. VidGen T2V handles complex scene descriptions that include people, motion, and background context, consistently producing believable body structure across subjects.
Scene extensions
Extend any VidGen T2V or I2V output from 5 seconds up to 20 seconds in a single follow-up request. Video Extend is aware of the original clip's characters, body positions, and scene context — the continuation reads as one shot rather than a stitch.
Creator content pipelines
Embed VidGen I2V directly into a creator-facing product — let users upload a portrait or describe a scene to generate personalized, repeatable social video on demand. The integration pattern is a single task endpoint with polling; hand off the task ID to the client for async UX if you prefer not to block on the server.
Human-optimized output
Accurate body structure and portrait detail across text, image, and extension inputs — VidGen's core advantage over general-purpose video models.
Built for action and motion
Handles complex scene composition, dynamic movement, and background audio — strong where general-purpose models lose consistency.
Three models, one request pattern
T2V, I2V, and Video Extend share the same prompt, media, and callback structure — switch modes without changing your integration.