# Vidu Q1 Reference Guide: Consistent Characters and Products from Up to 7 Images

Canonical page: https://fuser.studio/articles/vidu-reference-guide

How to use Vidu Q1 Reference: pick and prepare up to 7 reference images, write prompts that name each subject, set movement amplitude, and what our real test clips kept or lost.

[All guides](https://fuser.studio/articles) · [Vidu Reference in Fuser](https://fuser.studio/models/vidu-reference)

**Quick answer:** Vidu Q1 Reference makes a five-second, 1080p clip from a text prompt plus 1 to 7 reference images, and keeps the people, products and places in those images recognisable in the video ([Vidu API docs](https://platform.vidu.com/docs/reference-to-video)). Use one clean image per subject, describe each subject in the prompt by what is visible in its image, and set **Movement Amplitude** to large when the subject has to travel through the frame. In our test, a character, a product and a location went in as three separate images and came out as one coherent shot with the character's face, hair and jacket intact; small label text was the first thing to break.

## What Vidu Q1 Reference is

Most image-to-video models animate one picture: your image becomes the first frame. A reference-to-video model works differently. It reads several images as a cast list and builds a new scene around them, so the character from one image can hold the product from another inside the location from a third. None of the reference images has to be the opening frame.

Fuser ships this as the [Vidu Reference node](https://docs.fuser.studio/docs/nodes/video/vidu-reference), which calls the Vidu Q1 reference-to-video model ([Vidu API docs](https://platform.vidu.com/docs/reference-to-video)). The node's **Model** dropdown defaults to Q1. It also lists an older "Vidu" option that accepts only three references; there is no reason to pick it for new work.

What the model accepts and returns:

- **References:** 1 to 7 images in PNG, JPEG or WebP, at least 128 × 128 px and no more stretched than 4:1 in either direction ([Vidu API docs](https://platform.vidu.com/docs/reference-to-video)). Fuser rejects an eighth image before the job is sent.
- **Prompt:** up to 1,500 characters in Fuser.
- **Output:** a fixed five-second clip at 1080p. Our three clips each came back as 1920 × 1080 at 24 fps.
- **Aspect ratio:** 16:9, 9:16 or 1:1.
- **Movement Amplitude:** auto, small, medium or large.
- **Seed:** 0 to 65,535 in Fuser, randomised on each run unless you set it.
- **Add Music:** off by default. Switch it on for a generated background track; there is no speech or sound design.

## Test 1: a character, a product and a place

We reused assets from earlier guides: the courier character from [our consistent-characters guide](https://fuser.studio/articles/consistent-characters-ai) (silver asymmetric bob, mustard rain jacket with two reflective stripes, teal messenger bag), a packshot of a SOLA hand wash bottle, and an empty bakery counter. All three went into one Vidu Q1 Reference node with this prompt: "The woman in the yellow rain jacket stands behind the wooden bakery counter in warm morning light. She picks up the SOLA hand wash bottle, turns it so the label faces the camera, and smiles. Medium shot, slow push-in."

![Character, product and bakery reference images beside the Vidu Q1 Reference clip: the woman picks up the SOLA bottle and turns its label to camera.](https://statics.fuser.studio/cms/5cdf5547-67de-4d70-bc78-f9536261f6b3)

_Three references in, one five-second clip out. One run, movement amplitude auto, seed 4242. Generated 28 September 2026._

What held: her face, the bob, the jacket's two stripes and the teal strap stayed consistent for all five seconds. She did every action in order: reach, lift, turn the label to camera, smile. The camera pushed in slowly as asked. The SOLA wordmark and the label's sage-green panel matched the packshot.

What changed: the bakery was rebuilt, not copied. Loaves, croissants, the chalkboard sign and a coffee cup carried over, but the room gained a curtained window, a hanging lamp and new shelving, and the counter faced the camera head-on. A ring appeared on one of her fingers that is not in the reference. And the small lines of text on the label turned into letter-like shapes:

![Close crop of the last frame: the SOLA wordmark is intact while the small lines of label text underneath have turned into unreadable letter shapes.](https://statics.fuser.studio/cms/863265d4-8bc5-4834-9b96-ac35a50e9527)

_Last frame, cropped. The wordmark survives; the fine print does not._

So treat reference-to-video as "these things, in a scene like this", not as compositing. If a label has to be legible in the final ad, put the real packshot over the frame afterwards, or keep the product small in the shot and cut to a clean product still.

## Test 2: what movement amplitude really does

For the second test we used the character image alone and asked her to walk toward the camera on a wet neon street, glance left, then look back. We ran the same prompt and seed twice, changing only **Movement Amplitude**.

![Same reference, prompt and seed run twice in Vidu Q1 Reference: movement amplitude auto (left) and large (right), a woman in a mustard jacket on a wet neon street.](https://statics.fuser.studio/cms/3c77d62c-fd0d-4345-bf06-15743e8a65df)

_One reference image, identical prompt and seed; only movement amplitude changed. Generated 28 September 2026._

- **Auto:** she barely moved forward. She did turn her head to the left and back, and the camera stayed nearly still.
- **Large:** she walked steadily toward the camera and the street moved past her, but the head turn mostly disappeared.

Two lessons. First, if the subject has to travel, set the amplitude yourself; auto can read "walk" as a gentle drift. Second, the seed does not lock the scene when another setting changes: the two street clips have different shopfronts and lighting. In both runs she kept her hands in her pockets, the pose from the reference image, even though walking would usually swing the arms. The reference pose carries weight, so choose one close to the action you want.

## How to prepare reference images

- **One subject per image.** A character, a product or a location, each in its own image. Vidu's own API lets you name each subject and point to it in the prompt as @name ([Vidu API docs](https://platform.vidu.com/docs/reference-to-video)); Fuser's node takes a plain list of images, so the prompt is the only way to tell the model which image is which.
- **Plain backgrounds for people and products.** Our character and bottle were shot on flat grey, so nothing from their backgrounds leaked into the bakery.
- **Show the details that must survive.** The model kept the jacket stripes and the strap because they were large and clear in the reference. It lost the fine print because it was small.
- **Pick the pose you need.** See test 2: the reference pose carried into both clips.
- **Keep the set small.** Seven is the ceiling, not a target. Every extra image is another thing the prompt has to place.

## How to write the prompt

1. **Name each subject by what is visible in its image.** "The woman in the yellow rain jacket", "the SOLA hand wash bottle". Use the same words every time you run the shot.
2. **Put the location in words even when you supply an image of it.** "Behind the wooden bakery counter in warm morning light" gave the model a place to stand her.
3. **Write actions as a short sequence.** "Picks up, turns it so the label faces the camera, and smiles" was followed in order. Five seconds fits two or three beats.
4. **Add one camera instruction.** "Medium shot, slow push-in" was honoured. Stacking several moves in five seconds leaves little room for any of them.
5. **Leave appearance to the images.** Describe clothing only as far as you need to point at the right reference, not to restyle it.

## Build it as a workflow in Fuser

Every reference is its own node, so the same character image can feed a still-image edit, this video node and the next shot at the same time. A typical chain: create the character with [GPT Image](https://fuser.studio/models/gpt-image) or [Gemini Image](https://fuser.studio/models/gemini-image), connect the character, product and location images to the Vidu Reference node's **Images** input, then send the clip to [SeedVR upscale](https://fuser.studio/models/seedvr-upscale) or add sound with [Mirelo SFX](https://fuser.studio/models/mirelo-sfx). Swap the product image and rerun, and the rest of the graph stays as it was.

For single-image animation, where one exact frame must open the shot, a start-frame model such as [Kling 3.0](https://fuser.studio/models/kling-3-0-video) is the better fit, and the [first and last frame guide](https://fuser.studio/articles/first-last-frame-video-guide) covers fixing both ends of a clip. [Veo 3.1](https://fuser.studio/models/veo) and [Seedance 2](https://fuser.studio/models/seedance-2) also take multiple references; see [the best image-to-video models](https://fuser.studio/articles/best-image-to-video-models) to compare them.

## Vidu Q1 Reference settings in Fuser.

What each control does and what to start with.

### Vidu Reference node

| Setting | What it does | Start with |
| --- | --- | --- |
| Images | 1 to 7 reference images: characters, products, locations. | One clean image per subject |
| Prompt | Scene, action and camera, up to 1,500 characters. | Name each subject by what is visible |
| Model | Q1, or the older Vidu option (3 images max). | Q1 (default) |
| Movement Amplitude | How much the subject and camera move. | Large for walking or travel; auto for small gestures |
| Aspect Ratio | 16:9, 9:16 or 1:1. | Match the platform you are posting to |
| Seed | 0 to 65,535; same inputs and seed repeat a result. | Fix it while you tune the prompt |
| Add Music | Adds a generated background track. | Off; add sound in a later step |

## Questions, answered.

### How many reference images does Vidu Q1 accept?

Up to seven. In Fuser, connect them all to the Images input of the Vidu Reference node with the Q1 model selected; the older Vidu option is limited to three.

### How long are Vidu Q1 Reference videos?

Five seconds at 1080p. Duration and resolution are fixed for this model; you choose the aspect ratio (16:9, 9:16 or 1:1).

### Does Vidu reference-to-video keep product text readable?

Large, simple lettering usually survives; small text does not. In our test the SOLA wordmark held while the smaller label lines turned into unreadable shapes. Overlay the real packshot afterwards when the label must be legible.

### What is the difference between reference-to-video and image-to-video?

Image-to-video starts from your image as the first frame. Reference-to-video treats your images as subjects to place in a new scene, so none of them has to be the opening frame and several can appear together.

### Does Vidu Q1 Reference generate audio?

Only optional background music, which is off by default in Fuser. It does not generate speech or sound effects; add those with a separate audio node.

## Put your cast in one shot.

Wire your character, product and location into Vidu Q1 Reference on one canvas.

[Open Vidu Reference in Fuser](https://fuser.studio/models/vidu-reference) · [Explore all guides](https://fuser.studio/articles)

## More articles

- [Wan 2.6 Guide: Multi-Shot Video, Audio and Reference Clips](https://fuser.studio/articles/wan-2-6-guide.md)
- [Luma Ray 3.2 Guide: Prompts, Keyframes, Loops and Reframe](https://fuser.studio/articles/luma-ray-3-2-guide.md)
- [How to Keep AI Characters Consistent Across Images and Video](https://fuser.studio/articles/consistent-characters-ai.md)
