Wan Collection

Wan × VidModel

Wan is Alibaba's open-source video generation family — Apache 2.0 licensed, DiT + Flow Matching architecture, and 86%+ VBench scores that outrank Sora. Eight models cover image-to-video, text-to-video, reference-to-video, and instruction-based video editing across Wan 2.6 and 2.7.

Model Group

Eight generation modes across two model generations, led by an open-source DiT architecture

Wan covers every core video generation workflow — animate still images, generate from text prompts, preserve subject identity from reference videos, or apply instruction-based edits to existing clips. Wan 2.7 extends 2.6 with 9-grid multi-angle input, up to 5 reference videos, first+last frame control, and a professional Pro Edit tier.

8

included models

15s

max video length

86%+

VBench score

Included models

Two generations, five input modes — pick by task type and generation target. All share the same API request pattern.

Wan 2.7 Image to Video

Alibaba's flagship I2V model — up to 15s at 1080p with 9-grid multi-angle input, first+last frame guidance, and physics-accurate motion from a single source image.

IMAGE INPUTStandard
Try it

Wan 2.7 Text to Video

Alibaba's 2.7 T2V model — up to 15s at 1080p from text prompts with multi-shot storytelling, automatic prompt expansion, and multi-character interaction support.

TEXT INPUTStandard
Try it

Wan 2.7 Reference to Video

Reference-driven video generation accepting up to 5 reference videos — preserves character identity, visual style, and motion patterns across up to 15s of 1080p output.

REFERENCE INPUTStandard
Try it

Wan 2.7 Edit

Instruction-based video editing — describe the change in plain text and Wan 2.7 Edit modifies the video while preserving the original motion, timing, and temporal structure.

AI INPUTStandard
Try it

Wan 2.7 Pro Edit

Professional-tier image editing with instruction-guided precision — 20+ points above standard Wan 2.7 Edit on benchmarks, supporting up to 9 reference images and 12-language prompts.

AI INPUTStandard
Try it

Wan 2.6 Image to Video

Proven open-source I2V foundation — up to 15s at 1080p from a single source image, with broad aspect ratio support, optional audio sync, and reproducible generation via seed.

IMAGE INPUTStandard
Try it

Wan 2.6 Text to Video

Scene-aware T2V with multi-shot segmentation enabled by default — up to 15s at 1080p from text, with native audio sync and character consistency across narrative cuts.

TEXT INPUTStandard
Try it

Wan 2.6 Reference to Video

Proven reference-to-video foundation — up to 3 reference videos, 5 or 10s output at 1080p, with character identity and visual style preserved across generated clips.

REFERENCE INPUTStandard
Try it

Pick the right Wan model

Eight models across two generations — pick by input type, generation goal, and required output duration.

ModelInputDurationBest For
Source imageUp to 15sMulti-angle animation, first+last frame control
Text promptUp to 15sMulti-shot narratives, complex scene composition
Reference video ×5Up to 15sCharacter consistency, identity-preserving generation
Video + instructionUp to 15sNon-destructive video editing, style transfer
Image + instructionProfessional image editing, multi-reference fusion
Source imageUp to 15sProven I2V foundation, broad resolution support
Text promptUp to 15sScene-aware multi-shot T2V, audio sync
Reference video ×3Up to 10sEstablished character consistency baseline
PRODUCTION PLAYBOOKS

Match each Wan model to a real video workflow

I2V, T2V, R2V, and instruction-based editing — one API pattern handles all eight modes.

01
2.7 Image to Video

9-grid product animation

Animate a product or subject image into a 1080p clip with physics-accurate motion — cloth, reflections, and secondary movement rendered faithfully at 24fps. Wan 2.7 I2V's 9-grid input mode feeds up to 9 reference images for richer multi-angle scene understanding, and first+last frame control lets you define a precise start-to-end motion arc in a single request.

02
2.7 Text to Video

Multi-shot narrative generation

Generate up to 15 seconds of 1080p video from text prompts with multi-shot mode on by default — Wan 2.7 T2V segments the prompt into scenes and maintains character and environment consistency across every cut, without manual stitching. Built-in LLM prompt expansion enriches short instructions into richer generation directives automatically.

03
2.7 Reference to Video

Character-consistent multi-shot sequences

Wan 2.7 R2V accepts up to 5 reference videos simultaneously — tagged @Video1–@Video5 in the prompt — and generates new video clips that faithfully preserve each subject's identity, visual style, and motion dynamics. Three times the reference capacity of Wan 2.6 R2V, and up to 15 seconds of output versus 10.

04
2.7 Edit

Instruction-based video editing

Wan 2.7 Edit applies natural-language editing instructions to an existing video clip — swapping backgrounds, shifting lighting and mood, changing subject clothing, or applying style transfer — while preserving the original motion, timing, and temporal structure. No mask, no segmentation: describe the change and the model handles it.

WanAPI
1
2.7 I2V
wan-2.7-i2v
2
2.7 T2V
wan-2.7-t2v
3
2.7 R2V
wan-2.7-r2v
4
2.7 Edit
wan-2.7-edit

Open-source with 86%+ VBench

Wan 2.6 is Apache 2.0 licensed and open-weight — the same model that scores 86.22% on VBench, above Sora's 84.28%, freely available for commercial use, fine-tuning, and deployment.

Five generation modes in one API

Image, text, reference, video editing, and professional image editing all use the same task endpoint and polling pattern — swap generation modes without changing your integration.

Instruction editing without masks

Wan 2.7 Edit applies background swaps, lighting shifts, and style transfers from plain-text instructions alone — no segmentation masks, no extra tooling, no regeneration of the full clip.