
Wan × VidModel
Wan is Alibaba's open-source video generation family — Apache 2.0 licensed, DiT + Flow Matching architecture, and 86%+ VBench scores that outrank Sora. Eight models cover image-to-video, text-to-video, reference-to-video, and instruction-based video editing across Wan 2.6 and 2.7.
Eight generation modes across two model generations, led by an open-source DiT architecture
Wan covers every core video generation workflow — animate still images, generate from text prompts, preserve subject identity from reference videos, or apply instruction-based edits to existing clips. Wan 2.7 extends 2.6 with 9-grid multi-angle input, up to 5 reference videos, first+last frame control, and a professional Pro Edit tier.
8
included models
15s
max video length
86%+
VBench score
Included models
Two generations, five input modes — pick by task type and generation target. All share the same API request pattern.

Wan 2.7 Image to Video
Alibaba's flagship I2V model — up to 15s at 1080p with 9-grid multi-angle input, first+last frame guidance, and physics-accurate motion from a single source image.

Wan 2.7 Text to Video
Alibaba's 2.7 T2V model — up to 15s at 1080p from text prompts with multi-shot storytelling, automatic prompt expansion, and multi-character interaction support.

Wan 2.7 Reference to Video
Reference-driven video generation accepting up to 5 reference videos — preserves character identity, visual style, and motion patterns across up to 15s of 1080p output.

Wan 2.7 Edit
Instruction-based video editing — describe the change in plain text and Wan 2.7 Edit modifies the video while preserving the original motion, timing, and temporal structure.

Wan 2.7 Pro Edit
Professional-tier image editing with instruction-guided precision — 20+ points above standard Wan 2.7 Edit on benchmarks, supporting up to 9 reference images and 12-language prompts.

Wan 2.6 Image to Video
Proven open-source I2V foundation — up to 15s at 1080p from a single source image, with broad aspect ratio support, optional audio sync, and reproducible generation via seed.

Wan 2.6 Text to Video
Scene-aware T2V with multi-shot segmentation enabled by default — up to 15s at 1080p from text, with native audio sync and character consistency across narrative cuts.

Wan 2.6 Reference to Video
Proven reference-to-video foundation — up to 3 reference videos, 5 or 10s output at 1080p, with character identity and visual style preserved across generated clips.
Pick the right Wan model
Eight models across two generations — pick by input type, generation goal, and required output duration.
| Model | Input | Duration | Best For |
|---|---|---|---|
| Source image | Up to 15s | Multi-angle animation, first+last frame control | |
| Text prompt | Up to 15s | Multi-shot narratives, complex scene composition | |
| Reference video ×5 | Up to 15s | Character consistency, identity-preserving generation | |
| Video + instruction | Up to 15s | Non-destructive video editing, style transfer | |
| Image + instruction | — | Professional image editing, multi-reference fusion | |
| Source image | Up to 15s | Proven I2V foundation, broad resolution support | |
| Text prompt | Up to 15s | Scene-aware multi-shot T2V, audio sync | |
| Reference video ×3 | Up to 10s | Established character consistency baseline |
Match each Wan model to a real video workflow
I2V, T2V, R2V, and instruction-based editing — one API pattern handles all eight modes.
9-grid product animation
Animate a product or subject image into a 1080p clip with physics-accurate motion — cloth, reflections, and secondary movement rendered faithfully at 24fps. Wan 2.7 I2V's 9-grid input mode feeds up to 9 reference images for richer multi-angle scene understanding, and first+last frame control lets you define a precise start-to-end motion arc in a single request.
Multi-shot narrative generation
Generate up to 15 seconds of 1080p video from text prompts with multi-shot mode on by default — Wan 2.7 T2V segments the prompt into scenes and maintains character and environment consistency across every cut, without manual stitching. Built-in LLM prompt expansion enriches short instructions into richer generation directives automatically.
Character-consistent multi-shot sequences
Wan 2.7 R2V accepts up to 5 reference videos simultaneously — tagged @Video1–@Video5 in the prompt — and generates new video clips that faithfully preserve each subject's identity, visual style, and motion dynamics. Three times the reference capacity of Wan 2.6 R2V, and up to 15 seconds of output versus 10.
Instruction-based video editing
Wan 2.7 Edit applies natural-language editing instructions to an existing video clip — swapping backgrounds, shifting lighting and mood, changing subject clothing, or applying style transfer — while preserving the original motion, timing, and temporal structure. No mask, no segmentation: describe the change and the model handles it.
Open-source with 86%+ VBench
Wan 2.6 is Apache 2.0 licensed and open-weight — the same model that scores 86.22% on VBench, above Sora's 84.28%, freely available for commercial use, fine-tuning, and deployment.
Five generation modes in one API
Image, text, reference, video editing, and professional image editing all use the same task endpoint and polling pattern — swap generation modes without changing your integration.
Instruction editing without masks
Wan 2.7 Edit applies background swaps, lighting shifts, and style transfers from plain-text instructions alone — no segmentation masks, no extra tooling, no regeneration of the full clip.