VidGen Collection

VidGen × VidModel

VidGen is built around human subjects — delivering exceptional body structure and portrait accuracy across complex action scenes, with background audio support. Access T2V, I2V, and Video Extend through VidModel's unified API.

Model Group

Exceptional human body and portrait rendering, across three generation modes

VidGen is optimized for human subjects — producing accurate body structure, natural movement, and facial detail even in complex action scenes. Choose from text-to-video, image-to-video, or video extension, all on the same API pattern.

3

included models

20s

max video length

Text + Image

input modes

Included models

Pick the model by input type and output goal. Every card maps to a production API workflow.

VidGen 1.0 Image to Video

Animate images into video with accurate human body structure and natural movement — strong in complex action scenes with background audio support.

IMAGE INPUTStandard
Try it

VidGen 1.0 Video Extend

Extend VidGen-generated video up to 20 seconds while preserving body structure, character consistency, and scene continuity.

VIDEO INPUTUtility
Try it

VidGen 1.0 Text to Video

Generate video from text prompts with exceptional human body structure and portrait accuracy — ideal for complex action scenes and audio-rich content.

TEXT INPUTStandard
Try it

Pick the right model

All three models share the same request pattern — pick by input type and generation goal.

ModelInputMax OutputBest For
Portrait image20sSpokesperson videos, portrait animation
Text prompt20sAd campaigns, scripted character scenes
Existing video+20sScene extension, continuations
PRODUCTION PLAYBOOKS

Put VidGen's human-optimized output to work

Three models, one API pattern — pick by input type and generation goal.

01
Image to Video

Spokesperson videos

Animate a portrait into a talking, moving character for product pitches, feature walkthroughs, or brand introductions. VidGen's human-body specialization means the output holds accurate posture, limb placement, and natural facial expression through the full clip — without the drift typical of general-purpose video models.

02
Text to Video

Human-led ad campaigns

Generate people-centered video concepts from text prompts at speed — rapid variant creation without actors, scheduling, or a production shoot. VidGen T2V handles complex scene descriptions that include people, motion, and background context, consistently producing believable body structure across subjects.

03
Video Extend

Scene extensions

Extend any VidGen T2V or I2V output from 5 seconds up to 20 seconds in a single follow-up request. Video Extend is aware of the original clip's characters, body positions, and scene context — the continuation reads as one shot rather than a stitch.

04
Image to Video

Creator content pipelines

Embed VidGen I2V directly into a creator-facing product — let users upload a portrait or describe a scene to generate personalized, repeatable social video on demand. The integration pattern is a single task endpoint with polling; hand off the task ID to the client for async UX if you prefer not to block on the server.

VidGenAPI
1
Image to Video
vidgen1.0-i2v
2
Text to Video
vidgen1.0-t2v
3
Video Extend
vidgen1.0-video-extend

Human-optimized output

Accurate body structure and portrait detail across text, image, and extension inputs — VidGen's core advantage over general-purpose video models.

Built for action and motion

Handles complex scene composition, dynamic movement, and background audio — strong where general-purpose models lose consistency.

Three models, one request pattern

T2V, I2V, and Video Extend share the same prompt, media, and callback structure — switch modes without changing your integration.