Wan-2.2

byTongyi Lab, Alibaba Group

Orchestrate cinematic motion, complex camera sweeps, and physics-accurate scene transitions with open-weight diffusion

Wan-2.2

How Wan-2.2 Video works

From single text prompts to dual-frame interpolation, direct every dimension of camera movement and physical motion.

Frame your prompt and visual assets

Frame your prompt and visual assets

Enter a cinematic prompt or upload anchor frames, setting initial and terminal images for precise frame interpolation.

Configure motion dynamics and model tier

Configure motion dynamics and model tier

Select between the agile Lite 5B model or the heavy-duty Pro 14B model, adjusting shift values from 3.5 to 7.0 for tailored motion pacing.

Render volumetric video sequences

Render volumetric video sequences

Generate smooth 81 to 121 frame video clips at up to 720p with stable camera motion and realistic physical behavior.

What Wan-2.2 Video is good at

An MoE diffusion pipeline engineered for volumetric lighting, physical consistency, and granular motion control.

Mixture-of-Experts Physics Modeling

Mixture-of-Experts Physics Modeling

Harnessing 14 billion active parameters per step from a 27 billion pool, the model separates structural layout from fine textural refinement to render cohesive real-world gravity, fluid dynamics, and stable camera sweeps.

First-to-Last Frame Interpolation

First-to-Last Frame Interpolation

The Pro (14B) tier natively connects starting and ending keyframes with seamless visual logic, eliminating morphing artifacts while maintaining subject consistency and lighting transitions across 81 to 121 frames.

Video-to-Video Motion Transfer

Video-to-Video Motion Transfer

Transform existing video footage into stylized visual aesthetics or alter subject identities while retaining camera choreography and organic temporal pacing across the sequence.

Turbo and Distilled Rapid Iteration

Turbo and Distilled Rapid Iteration

Switch to dedicated Turbo and Distill inference modes for fast concept verification, dramatically cutting render times without sacrificing compositional stability before final high-step rendering.

Made with Wan-2.2 Video

Explore high-fidelity video generations spanning complex physics, dynamic camera sweeps, and rich atmospheric environments.

High-speed automotive drift with physically consistent terrain interaction

High-speed automotive drift with physically consistent terrain interaction

Zero-gravity physical interaction and volumetric spatial depth

Zero-gravity physical interaction and volumetric spatial depth

Choreographed fabric dynamics and high-speed motion capture

Choreographed fabric dynamics and high-speed motion capture

Dynamic fluid simulation and steady panoramic camera pullback

Dynamic fluid simulation and steady panoramic camera pullback

Intricate mechanical simulation and macro-scale camera control

Intricate mechanical simulation and macro-scale camera control

What people build with Wan-2.2 Video

How directors, 3D animators, and commercial studios use cinematic video synthesis across production workflows.

Cinematic Film & Previs Directors

01

Choreograph complex shot sequences with precise camera maneuvers—including crash zooms, sweeping dollies, and panoramic tilts—grounded in realistic gravity and lighting.

High-Fashion Editorial Animators

02

Bring lookbooks and editorial stills to life with natural fabric movement, believable hair dynamics, and nuanced character gestures without warping silhouettes.

Commercial Product Animators

03

Render photorealistic product commercial sequences, highlighting material textures like brushed metals, frosted glass, and flowing liquids with production-grade lighting.

VFX & Video Restyling Artists

04

Apply complete aesthetic overhauls to live-action footage using video-to-video pipelines, re-skinning characters and environments while locking the original camera track.

Archival & Documentarian Storytellers

05

Interpolate historical imagery into fluid motion using First-and-Last Frame conditioning, restoring archival photographic moments into living historical vignettes.

Frequently Asked Questions

The Pro (14B) model utilizes a 27-billion total parameter MoE architecture with 14 billion active parameters per step, delivering superior physical simulation, multi-subject adherence, complex camera choreographies, and native support for First-to-Last Frame (FLF2V) interpolation. The Lite (5B) model is an agile, budget-friendly tier optimized for rapid text-to-video generation of shorter clips and standard image animation where extreme multi-object interaction is not needed.

Use the Turbo and Distill variants when you need rapid feedback for creative drafting, layout testing, and interactive storyboarding. These modes leverage step-distillation to speed up generation significantly at reduced computational cost, with only slight tradeoffs in fine background textural detail compared to the standard full-step inference pipelines.

First-and-Last Frame conditioning lets you provide both a starting image and an ending image to the Pro (14B) model. The diffusion pipeline then generates a logically coherent visual progression between the two frames across 81 to 121 frames, maintaining subject identity, lighting continuity, and natural physical motion throughout the transition.

Structure prompts using the sequence: Subject + Scene + Motion + Camera Movements + Visual Style. Separate chronological events with periods instead of commas, as periods cue the model to sequence actions over time rather than rendering them simultaneously. To calibrate movement intensity, adjust the shift parameter between 3.5 for gentle portrait motion and up to 7.0 for fast-paced kinetic action.

The model is not designed for generating legible text overlays, high-speed human combat with chaotic limb crossing, or extreme abstract art that violates basic physical laws. Because its MoE pipeline is heavily trained on realistic spatial logic and volumetric consistency, scenes requiring impossible non-Euclidean geometry may produce visual artifacts.

Try Wan-2.2 Video on Fuser

Orchestrate cinematic motion, complex camera sweeps, and physics-accurate scene transitions with open-weight diffusion