Fuser Apps are here 🚀
Fuser Apps are here. Free generations for the next month 💫
Let's GobyTongyi Lab, Alibaba Group
Orchestrate cinematic motion, complex camera sweeps, and physics-accurate scene transitions with open-weight diffusion
From single text prompts to dual-frame interpolation, direct every dimension of camera movement and physical motion.
An MoE diffusion pipeline engineered for volumetric lighting, physical consistency, and granular motion control.
Explore high-fidelity video generations spanning complex physics, dynamic camera sweeps, and rich atmospheric environments.
How directors, 3D animators, and commercial studios use cinematic video synthesis across production workflows.
The Pro (14B) model utilizes a 27-billion total parameter MoE architecture with 14 billion active parameters per step, delivering superior physical simulation, multi-subject adherence, complex camera choreographies, and native support for First-to-Last Frame (FLF2V) interpolation. The Lite (5B) model is an agile, budget-friendly tier optimized for rapid text-to-video generation of shorter clips and standard image animation where extreme multi-object interaction is not needed.
Use the Turbo and Distill variants when you need rapid feedback for creative drafting, layout testing, and interactive storyboarding. These modes leverage step-distillation to speed up generation significantly at reduced computational cost, with only slight tradeoffs in fine background textural detail compared to the standard full-step inference pipelines.
First-and-Last Frame conditioning lets you provide both a starting image and an ending image to the Pro (14B) model. The diffusion pipeline then generates a logically coherent visual progression between the two frames across 81 to 121 frames, maintaining subject identity, lighting continuity, and natural physical motion throughout the transition.
Structure prompts using the sequence: Subject + Scene + Motion + Camera Movements + Visual Style. Separate chronological events with periods instead of commas, as periods cue the model to sequence actions over time rather than rendering them simultaneously. To calibrate movement intensity, adjust the shift parameter between 3.5 for gentle portrait motion and up to 7.0 for fast-paced kinetic action.
The model is not designed for generating legible text overlays, high-speed human combat with chaotic limb crossing, or extreme abstract art that violates basic physical laws. Because its MoE pipeline is heavily trained on realistic spatial logic and volumetric consistency, scenes requiring impossible non-Euclidean geometry may produce visual artifacts.
Orchestrate cinematic motion, complex camera sweeps, and physics-accurate scene transitions with open-weight diffusion