
VidVox × VidModel
VidVox delivers fast generation and lifelike visual realism — produce character video, blend reference images, or transfer motion across subjects. Access all VidVox models through VidModel's unified API.
Fast generation and lifelike realism, across four character video modes
VidVox is built for speed and visual realism. Pick the model by input type — portrait image, reference frames, or text prompt — and generate high-quality character video using the same API pattern across every task.
4
included models
15s
max video output
Image + Text + Ref
input modes
Included models
Each model handles a distinct generation task. Pick by input type and the output speed you need.

VidVox 1.0 Reference to Video
Blend multiple images into a video, replace characters, or change actions in video.

VidVox 1.0 Image to Video
Animate a single portrait image into realistic talking-character video — up to 15s HD per generation, with fast output and lifelike results.

VidVox 1.0 Flash Image to Video
The fastest VidVox variant — generate talking-character video with lifelike realism at significantly reduced generation time.
Pick the right model
Four models, one API pattern — pick by input type and the speed you need.
| Model | Input | Speed | Best For |
|---|---|---|---|
| Portrait photo | Standard | Talking personas, product faces | |
| Reference images | Standard | Multi-character scenes, motion transfer | |
| Portrait batch | Fast | High-volume pipelines, personalization | |
— | Text prompt | Standard | Script-to-character, no photo needed |
Match each VidVox model to a real character workflow
Four models, one API pattern — pick by input type and the speed you need.
Talking product personas
Animate a brand face or spokesperson portrait into lifelike talking-character video for product pitches, feature highlights, or brand introductions. VidVox I2V is built for speed and visual realism — the output character moves expressively rather than rigidly, with lip movement that reads as natural at HD resolution.
Multi-reference character scenes
VidVox R2V accepts multiple reference images and three distinct modes of operation: blend separate character photos into a shared scene, extract and transfer motion from a reference clip, or swap a subject into existing video. One API call handles all three — switch between modes by passing referenceImages or referenceVideo.
High-volume content batches
VidVox Flash cuts generation time significantly compared to the standard I2V tier while preserving the same character quality and lifelike output. Kick off tasks concurrently with Promise.all and poll them in parallel — the only practical limit is your account rate cap, not sequential generation time.
Script-to-character video
Describe a character and scene in a prompt to generate talking-character video with no portrait or reference image required. VidVox T2V is the fastest way to prototype a character — define appearance, personality, and movement in a prompt brief, and use the output to validate the concept before sourcing real portrait photos.
Built for speed
VidVox generates output significantly faster than general-purpose video models — built for rapid iteration and high-volume production pipelines.
Multi-modal inputs
Drive video from portrait images, reference clips, or text prompts — one API pattern handles every VidVox generation type.
15-second HD output
Generate lifelike HD character video in one request — fast output with high visual fidelity, ready for production delivery.