Fuser Apps are here 🚀
Fuser Apps are here. Free generations for the next month 💫
Let's GobyShengShu Technology
Lock character identities across sequential scenes and blend outfits, props, and cinematic sets from multiple visual references
Harness multi-image visual feature fusion to anchor faces, garments, and environments across cinematic four-second shots.
Pin character identities, fuse distinct visual assets, and direct smooth camera movements without identity drift.
Explore how visual reference fusion unites independent character, wardrobe, and environment stills into cohesive video scenes.
From episodic animation storyboards to high-fashion lookbooks, discover how creators maintain strict character fidelity across shots.
Vidu Q1 is the newer generation model that supports native background music generation and provides superior fidelity when fusing multiple complex references. In contrast, the original Vidu model offers classic motion dynamics and does not support the background music toggle, making it primarily useful for legacy project continuity or fine-tuning comparisons.
Use Vidu Q1 for almost all multi-reference workflows, particularly when you need synchronized atmospheric audio or need to fuse distinct character, clothing, and background images cleanly. Reach for the original Vidu model only when you want to match the specific motion cadence of earlier generations or benchmark complex shot behavior.
You can upload up to seven reference images into the model. This multi-entity fusion allows you to combine separate images for character faces, distinct outfits, props, and custom background environments into a single cohesive scene.
The movement amplitude parameter controls the physical speed and camera dynamics across four settings: Auto, Small, Medium, and Large. For character portraits and close-up facial shots, use Small or Medium to avoid facial morphing and limb distortion; reserve Large for wide-angle landscape dynamics like ocean swells or heavy winds.
Avoid using Vidu Reference for text-only generation without reference images, as the underlying architecture relies heavily on visual anchors. Additionally, avoid hyper-kinetic action shots, complex acrobatics, or fine finger manipulations such as typing on a keyboard or playing musical instruments, as rapid limb movements can cause visual melting artifacts.
Lock character identities across sequential scenes and blend outfits, props, and cinematic sets from multiple visual references