
Kling × VidModel
Kling is Kuaishou's video generation engine — physics-accurate motion, native multilingual audio, multi-shot storyboarding, and up to 1080p output. Three tiers cover every workflow from rapid iteration to cinematic production.
Physics-accurate video across four generation modes, with native audio and multi-shot control
Kling's model family covers image-to-video at Pro and Standard tiers, text-to-video with multi-shot storyboarding, and a speed-optimized Turbo tier for high-volume pipelines. All models share the same API pattern — switch tiers without changing your integration.
4
included models
15s
max video length
Native
audio generation
Included models
Pick by input type and quality target. Every card maps to a direct API workflow.

Kling v3-pro Image to Video
Kuaishou's flagship I2V model — up to 15s at 1080p with physics-aware motion, optional start/end frame guidance, and native audio generation.

Kling v3 Image to Video
Cost-effective Kling V3 image-to-video — structural fidelity to the source image, native audio, and up to 15s output at 1080p.

Kling v3 Text to Video
Physics-accurate text-to-video with multi-shot generation up to 15s, native multilingual audio, and up to 1080p — no source image required.

Kling v2.5-turbo Image to Video
Speed-optimized Kling I2V for high-throughput pipelines — fast generation, 720p output, 5 or 10s clips, and first/last frame control at reduced cost.
Pick the right model
Four tiers, one API pattern — pick by input type, quality target, and throughput need.
| Model | Input | Speed | Best For |
|---|---|---|---|
| Source image | Standard | Cinematic animation, production-quality output | |
| Source image | Standard | Structural fidelity, cost-efficient 1080p | |
| Text prompt | Standard | Multi-shot narratives, no image required | |
| Source image | Fast | High-volume pipelines, rapid iteration |
Match each Kling model to a real video workflow
Pro quality, cost-effective Standard, text-only T2V, or fast Turbo — one API pattern handles all four.
Cinematic portrait animation
Animate a source image into a cinematic 1080p clip with physics-accurate motion, optional start-to-end frame guidance for smooth transitions, and native audio generation — all in one request. Kling v3-pro's Omni One architecture with Visual Chain-of-Thought reasoning produces refined detail preservation and intricate motion that general-purpose I2V models miss.
Structural image animation at scale
Kling v3 Standard I2V prioritizes structural fidelity to the source image — the output holds the original composition, subject placement, and spatial relationships accurately across frames. The right tier when consistent structural adherence matters more than intricate motion detail, and when cost efficiency is a priority over Pro-tier quality.
Multi-shot scene generation
Generate full narrative video sequences from a text description — no source image required. Kling v3 T2V's multi-shot system produces up to 5 coherent scenes in a single API call (total ≤ 15s), each with its own prompt and duration, powered by Visual Chain-of-Thought reasoning that pre-plans physics, camera angles, and lighting before rendering.
High-throughput image animation
Kling v2.5-turbo is the speed-optimized I2V tier — designed for rapid iteration and high-volume pipelines where throughput and cost matter more than maximum resolution. At roughly 3× faster than Standard tiers and 25% lower cost than v2.1 Standard, it's the right model for personalization flows, preview generation, and any workflow that runs hundreds of requests per batch.
Physics-accurate motion
Omni One architecture with 3D Spacetime Joint Attention produces accurate cloth dynamics, gravity, hair, and secondary motion — the physics gap that separates Kling from general-purpose video models.
Native multilingual audio
Dialogue, ambient sound, and music generated alongside video in one request — English, Chinese, Japanese, Korean, and Spanish, with accurate lip-sync built in.
Multi-shot storyboarding
Define up to 5 narrative scenes with individual prompts and durations in a single T2V request — full video sequences without manual stitching.