Vidu Reference

byShengShu Technology

Lock character identities across sequential scenes and blend outfits, props, and cinematic sets from multiple visual references

Vidu Reference

How Vidu Reference works

Harness multi-image visual feature fusion to anchor faces, garments, and environments across cinematic four-second shots.

Anchor reference images

Anchor reference images

Upload up to seven visual references including character portraits, wardrobe stills, and environments.

Tag and script the action

Tag and script the action

Map reference elements with bracket tags like [R1] and specify camera direction alongside motion amplitude.

Generate cohesive video

Generate cohesive video

Synthesize a four-second cinematic clip preserving visual identity and optional synchronized score.

What Vidu Reference is good at

Pin character identities, fuse distinct visual assets, and direct smooth camera movements without identity drift.

Multi-Entity Concept Fusion

Multi-Entity Concept Fusion

Blend up to seven distinct reference images—including character faces, specific garments, and environment plates—into a unified composition.

Q1 Audio & Multi-Subject Engine

Q1 Audio & Multi-Subject Engine

Select the Q1 model tier for high-fidelity multi-subject tracking and synchronized atmospheric background music generation.

Identity-Locked Character Continuity

Identity-Locked Character Continuity

Prevent facial drift and structural distortion across consecutive video scenes by anchoring key identity markers directly from source portraits.

Dynamic Motion Amplitude Tuning

Dynamic Motion Amplitude Tuning

Control scene physics across small, medium, and large movement amplitudes to prevent facial warping while accommodating sweeping environmental movement.

Made with Vidu Reference

Explore how visual reference fusion unites independent character, wardrobe, and environment stills into cohesive video scenes.

Commercial luxury timepiece product reveal with custom liquid environment

Commercial luxury timepiece product reveal with custom liquid environment

Cinematic album visualizer maintaining character identity across weather dynamics

Cinematic album visualizer maintaining character identity across weather dynamics

Architectural interior rendering with imported bespoke furniture references

Architectural interior rendering with imported bespoke furniture references

Historical reportage recreation fusing archival costume and rugged geology

Historical reportage recreation fusing archival costume and rugged geology

High-contrast lookbook sequence blending garment textures and athletic movement

High-contrast lookbook sequence blending garment textures and athletic movement

What people build with Vidu Reference

From episodic animation storyboards to high-fashion lookbooks, discover how creators maintain strict character fidelity across shots.

Narrative Episodic Storyboarding

01

Maintain exact facial structures, wardrobe styles, and props across multi-shot storyboards without recurring character drift or identity deformation.

Fashion Lookbook & Garment Fusion

02

Superimpose real-world garment references onto stylized digital models to test drape, motion physics, and silhouette behavior in cinematic motion.

Brand Mascot Commercials

03

Bring 2D or 3D illustrated brand mascots into photorealistic or stylized live-action scenes while locking core character geometry and color palettes.

Music Video Visualizers

04

Generate atmospheric performance scenes synchronized with native background audio, combining character portraits with surreal abstract sets.

Archival History & Period Films

05

Combine vintage portrait photographs with period-accurate garments and locations to generate historically grounded video vignettes.

Frequently Asked Questions

Vidu Q1 is the newer generation model that supports native background music generation and provides superior fidelity when fusing multiple complex references. In contrast, the original Vidu model offers classic motion dynamics and does not support the background music toggle, making it primarily useful for legacy project continuity or fine-tuning comparisons.

Use Vidu Q1 for almost all multi-reference workflows, particularly when you need synchronized atmospheric audio or need to fuse distinct character, clothing, and background images cleanly. Reach for the original Vidu model only when you want to match the specific motion cadence of earlier generations or benchmark complex shot behavior.

You can upload up to seven reference images into the model. This multi-entity fusion allows you to combine separate images for character faces, distinct outfits, props, and custom background environments into a single cohesive scene.

The movement amplitude parameter controls the physical speed and camera dynamics across four settings: Auto, Small, Medium, and Large. For character portraits and close-up facial shots, use Small or Medium to avoid facial morphing and limb distortion; reserve Large for wide-angle landscape dynamics like ocean swells or heavy winds.

Avoid using Vidu Reference for text-only generation without reference images, as the underlying architecture relies heavily on visual anchors. Additionally, avoid hyper-kinetic action shots, complex acrobatics, or fine finger manipulations such as typing on a keyboard or playing musical instruments, as rapid limb movements can cause visual melting artifacts.

Try Vidu Reference on Fuser

Lock character identities across sequential scenes and blend outfits, props, and cinematic sets from multiple visual references