Vidu

byShengshu Technology

Expressive characters, clean cel-style anime and first-to-last-frame control, with spoken dialogue and ambience baked into the clip

How Vidu works

Start from a still or a single-action prompt, choose the model tier that fits the job, and draft small before you commit. The workflow rewards short, simple takes over crowded scenes.

Vidu workflow input

Frame the opening shot

Drop in a first-frame still, or write one dominant action plus a named camera move like a slow push-in.

Vidu workflow direction

Calibrate motion and duration

Pick a model tier, keep the clip short, set movement amplitude, and draft at 360p or 540p before committing.

Vidu workflow model and generated output

Render the take

Generate at up to 1080p; on Q3, quoted dialogue and ambience arrive synced in the same clip.

What Vidu is good at

Vidu is built around the character: believable faces, steady camera grammar, defined start and end frames, and anime motion that stays flat and clean. Q3 adds sound.

Subtle character micro-acting

Soft blinks, eye darts and lip movement stay on the character, with identity held from the first frame. Set movement amplitude to small for portraits; large can distort faces.

Turbo drafts, Pro finals

Iterate on stills with Q2 image-to-video Turbo, then switch to Q2 Pro for the best detail and expression fidelity. On Q3, the Turbo tier runs at roughly half the cost with the same feature set.

First-to-last frame control

Lock a start image and an end image, then describe the transition path so the middle doesn't warp. Keep both frames in the same style, lighting and framing for a convincing morph or reveal.

Clean anime and manga motion

Flat cel and 2D illustration motion where hair and cloth hold their shape instead of melting. Q1 text-to-video adds a dedicated anime style toggle; on other models, steer with the prompt or an illustrated first frame.

Made with Vidu

A spread of short takes, from documentary heat haze and product pours to neon choreography and macro mechanics, each built on one dominant motion and a named camera move.

Documentary capture with natural heat distortion and archival grain

Commercial product reveal with controlled fluid motion

Stylized performance frame with striking color contrast

High-contrast choreography under vivid neon light

Macro mechanical tracking with hard-surface stability

What people build with Vidu

Animators, filmmakers, ad teams and social creators reach for Vidu when the subject is a character or a controlled reveal, and the shot needs to read clearly in seconds.

Anime and manga short makers

01

Animate an illustrated keyframe into a clean 2D motion beat: a hair flick, a wind-caught scarf, a slow push-in. Cel-style output holds shape, so storytelling stays readable.

Character-driven short filmmakers

02

Build acting beats around a face: a blink, a glance away, a held breath. Vidu keeps the character consistent across the take and favors emotional close-ups over spectacle.

Product reveal campaign teams

03

Turn a hero still or a text brief into a stable pull-back or push-in reveal with soft rim light. Draft cheaply on a Turbo tier, then finalize on Pro.

Talking-character social creators

04

Q3 generates dialogue, effects and music together, so a quoted line lands synced to the character's lips. English, Japanese and Chinese are the officially supported dialogue languages.

Storyboard animatic artists

05

Bring panels to life with short single-action clips and explicit camera grammar. Use low resolutions and a fixed seed to compare takes before committing to a final render.

Versions
Q3, Q3 Turbo, Vidu Q2, Vidu Q1, Vidu
Inputs
Prompt (text), First Frame (image), Last Frame (image)
Output
video
Duration
1s, 2s, 3s, 9s, 10s, 11s, 12s, 13s, 14s, 15s, 16s, 4s, 5s, 6s, 7s, 8s
Resolution
360p, 540p, 720p, 1080p
Aspect ratios
4:3, 3:4, 16:9, 9:16, 1:1
Audio
Generates audio
Cost
48–3,395 credits per run

Vidu questions, answered

Q3 is the premium tier and Q3 Turbo is the faster, cheaper one with the same feature set, including native audio. Turbo costs roughly half as much, which makes it right for drafts and idea testing, while Q3 is the better choice for narrative fidelity in finals. Both support clips from 1 to 16 seconds and resolutions from 360p to 1080p.

Use Q2 for tight 4–8 second character shots with expressive faces and stable camera moves; use Q3 when you need native audio or clips longer than 8 seconds. Q2 offers image-to-video Turbo for fast drafts and Pro for final-quality detail. Q2 text-to-video has no end-frame input, and its turbo_mode only applies to image-to-video.

Q3 produces 1 to 16 second clips, but 4–5 seconds is the most stable range. Drift appears past roughly 8 seconds, so don't expect a single steady shot at 12–16 seconds. Q2 tops out at 8 seconds.

Yes, on Q3 and Q3 Turbo, with the Generate Audio toggle on. Dialogue, effects and music come out in one pass. Put spoken lines in quotes and name the ambience you want, such as rain or footsteps. English, Japanese and Chinese are the officially supported dialogue languages. Q2 offers only background music, and only at 4 seconds.

On Q1 text-only prompts, set Style to Anime. On other models and for image-driven runs the style toggle has no effect, so describe cel shading or 2D illustration in your prompt or start from an illustrated first frame, which animates especially cleanly.

It trails the best models on photoreal skin close-ups, complex physics, extreme action and readable on-screen text. Stacking several actions in one prompt also hurts results. It works best with one dominant motion and one camera move, on a character-led shot.

Try Vidu on Fuser

Expressive characters, clean cel-style anime and first-to-last-frame control, with spoken dialogue and ambience baked into the clip