VidVox Collection

VidVox × VidModel

VidVox delivers fast generation and lifelike visual realism — produce character video, blend reference images, or transfer motion across subjects. Access all VidVox models through VidModel's unified API.

Model Group

Fast generation and lifelike realism, across four character video modes

VidVox is built for speed and visual realism. Pick the model by input type — portrait image, reference frames, or text prompt — and generate high-quality character video using the same API pattern across every task.

4

included models

15s

max video output

Image + Text + Ref

input modes

Included models

Each model handles a distinct generation task. Pick by input type and the output speed you need.

VidVox 1.0 Reference to Video

Blend multiple images into a video, replace characters, or change actions in video.

REFERENCE INPUTStandard
Try it

VidVox 1.0 Image to Video

Animate a single portrait image into realistic talking-character video — up to 15s HD per generation, with fast output and lifelike results.

IMAGE INPUTStandard
Try it

VidVox 1.0 Flash Image to Video

The fastest VidVox variant — generate talking-character video with lifelike realism at significantly reduced generation time.

IMAGE INPUTStandard
Try it

Pick the right model

Four models, one API pattern — pick by input type and the speed you need.

ModelInputSpeedBest For
Portrait photoStandardTalking personas, product faces
Reference imagesStandardMulti-character scenes, motion transfer
Portrait batchFastHigh-volume pipelines, personalization
Text promptStandardScript-to-character, no photo needed
PRODUCTION PLAYBOOKS

Match each VidVox model to a real character workflow

Four models, one API pattern — pick by input type and the speed you need.

01
Image to Video

Talking product personas

Animate a brand face or spokesperson portrait into lifelike talking-character video for product pitches, feature highlights, or brand introductions. VidVox I2V is built for speed and visual realism — the output character moves expressively rather than rigidly, with lip movement that reads as natural at HD resolution.

02
Reference to Video

Multi-reference character scenes

VidVox R2V accepts multiple reference images and three distinct modes of operation: blend separate character photos into a shared scene, extract and transfer motion from a reference clip, or swap a subject into existing video. One API call handles all three — switch between modes by passing referenceImages or referenceVideo.

03
Flash I2V

High-volume content batches

VidVox Flash cuts generation time significantly compared to the standard I2V tier while preserving the same character quality and lifelike output. Kick off tasks concurrently with Promise.all and poll them in parallel — the only practical limit is your account rate cap, not sequential generation time.

04
Text to Video

Script-to-character video

Describe a character and scene in a prompt to generate talking-character video with no portrait or reference image required. VidVox T2V is the fastest way to prototype a character — define appearance, personality, and movement in a prompt brief, and use the output to validate the concept before sourcing real portrait photos.

VidVoxAPI
1
Image to Video
vidvox1.0-i2v
2
Flash I2V
vidvox1.0-i2v-flash
3
Reference to Video
vidvox1.0-r2v
4
Text to Video
vidvox1.0-t2v

Built for speed

VidVox generates output significantly faster than general-purpose video models — built for rapid iteration and high-volume production pipelines.

Multi-modal inputs

Drive video from portrait images, reference clips, or text prompts — one API pattern handles every VidVox generation type.

15-second HD output

Generate lifelike HD character video in one request — fast output with high visual fidelity, ready for production delivery.