Fuser Apps are here 🚀
Fuser Apps are here. Free generations for the next month 💫
Let's GobyAlibaba Cloud (Tongyi Lab)
Generate cinematic frames and multi-image compositions with true volumetric light, spatial depth, and bilingual English-Chinese prompt fidelity
From prompt composition to high-fidelity rendering in three focused creative steps.
Built on advanced multimodal Diffusion Transformer architecture for unmatched atmospheric light and compositional control.
A selection of cinematic stills, material studies, and narrative frames generated with precise lighting and spatial depth.
How visual storytellers, concept artists, and creative directors leverage spatial coherence in production.
Choose the text-to-image variant when generating completely novel visual concepts from pure natural language prompts up to 2,000 characters. Choose the image-to-image variant when you need to transform an existing image, execute style transfers, or composite elements while preserving core subject identity and spatial layout.
Use the image-to-image variant. It allows you to supply a reference image between 384px and 5000px up to 10MB, using textual cues to modify the environment or styling while retaining the character's facial structure and proportions across narrative frames.
The model is engineered for cinematic storyboards, multi-image narrative continuity, environmental concept art with realistic 3D perspective, and bilingual English-Chinese scene generation. Its Diffusion Transformer architecture excels at atmospheric volumetric lighting, natural depth of field, and rich textural realism without artificial gloss.
The model is not designed for fine typographic rendering, vector icons, or complex embedded text logos. It also operates under strict safety filters through hosted cloud endpoints, making it unsuitable for unfiltered or NSFW generation workflows.
The model was developed by Alibaba Cloud's Tongyi Lab (Qwen team) and released on December 17, 2025. Standard inference costs approximately $0.03 per image.
Generate cinematic frames and multi-image compositions with true volumetric light, spatial depth, and bilingual English-Chinese prompt fidelity