Veo

byGoogle DeepMind

Photoreal shots that look directed and sound finished, with dialogue, ambience and deliberate camera moves arriving together in one clip

How Veo works

Write the shot the way a director would, add keyframes or reference images if the look must hold, and choose a tier. Picture and sound come back together.

veo workflow input

Write the shot

Order your prompt like a director: lens and camera move, subject, action, light, then quoted dialogue and SFX lines, or start from a first-frame image.

veo workflow direction

Anchor look and motion

Add up to three Ingredients images to keep a character or prop recognizable, or set a last frame to define where the shot lands.

veo workflow model and generated output

Pick a tier and render

Draft on Fast or Lite, then re-run keepers on the full model with 8-second clips at 1080p or 4K.

What Veo is good at

Sound, continuity and control: spoken lines that sync, characters that survive across shots, keyframe reveals, and three tiers that match the stage of your project.

Dialogue and sound in one pass

Quoted lines come back lip-synced, with room tone and effects that match the scene. Name sounds explicitly with an SFX line so ambience is directed, not guessed.

Ingredients for recurring characters

Attach up to three reference images, covering a character, a prop and a location, and carry them into a new shot. Ingredients run at 8 seconds, so pair them with a steady camera instruction.

Three tiers, one workflow

Use Fast as the default workhorse for iterating prompts and dialogue timing, Lite for the cheapest high-volume drafts, and the full tier for final hero renders you have already proven out.

First and last frame reveals

Set an opening and closing image, then prompt the bridging move, like a slow arc or day turning to night. Choose end frames that are a believable journey within 4 to 8 seconds.

Made with Veo

Short, sound-rich scenes across craft, music, architecture, fashion and the deep sea, each prompt written with a lens, a camera move and an SFX line.

Craft-commercial frame with directional sound

Solo performance plate with acoustic resonance

Architectural light study with room tone

Vertical lookbook clip with a spoken cue

Underwater film still with aquatic ambience

What people build with Veo

Filmmakers, commercial directors, fashion teams, social creators and visualization studios reach for it when a clip has to look directed and sound finished.

Independent filmmakers

01

Block out a spoken scene or an establishing shot with lens, camera move and ambience in the prompt, and get a clip that feels directed. Each clip runs up to 8 seconds.

Commercial directors

02

Build product hero shots with matched ambience, prove them on Fast, then re-render the keepers on the full tier at 1080p or 4K for client-facing spots.

Fashion lookbook teams

03

Keep a garment and a model recognizable from clip to clip by feeding Ingredients references into locked 8-second shots.

Vertical social creators

04

Compose natively in 9:16 with sound already in the clip, so a spoken hook and a room-tone bed arrive without a separate edit.

Architectural visualization studios

05

Use first and last frames to travel from an exterior render into an interior one, with light and material shifts carried along the camera path.

Versions
Veo 3.1 Fast, Veo 3.1, Veo 3.1 Lite
Inputs
Prompt (text), First Frame (image), Last Frame (image), Ingredients (up to 3 images), Negative Prompt (text)
Output
video
Duration
4 seconds, 6 seconds, 8 seconds
Resolution
720p, 1080p, 4K
Aspect ratios
16:9 Horizontal, 9:16 Vertical
Audio
Generates audio
Cost
165–6,613 credits per run

Veo questions, answered

The full tier is for final hero renders, Fast is the default workhorse for iteration, and Lite is the most affordable option for drafts. The full tier gives the most fidelity but costs the most and renders slowest. Fast gives up some fidelity for speed and a lower cost. Lite costs least but has less polish, so check that settings like top resolutions or Ingredients behave as you need.

Draft on Fast or Lite, then re-run your keepers on the full model. Fast suits testing prompts, dialogue timing and keyframe setups. Lite suits high-volume previsualization and animatics. Reserve the full tier for high-resolution, client-facing shots you have already proven out, since it is not worth using for exploratory attempts.

Put the spoken line in quotation marks inside your prompt, for example: A woman says, "We have to leave now." Keep each line to about one breath so it fits the clip. Leave Generate Audio on, and add an SFX line for ambience. Keep to one speaker, since lip-sync and attribution drift when several characters talk.

Ingredients let you attach up to 3 reference images, such as a character, a prop and a location, to carry into a new shot. They only support 8-second output. Use clean, uncluttered references and put the clearest face first. Avoid combining them with a conflicting first frame, which muddies the result.

Clips run 4, 6 or 8 seconds at about 24 fps, in 16:9 or 9:16. Resolution options are 720p, 1080p and 4K, but 1080p, 4K and Ingredients only support 8 seconds. For anything longer than 8 seconds in a single take, this is the wrong tool.

Not reliably. On-screen text and subtitles tend to come out as garbled lettering, so add captions in post. You can also list terms like subtitles or text overlay in the negative prompt. Crowded multi-person dialogue and strict frame-level motion control are other weak spots.

Try Veo on Fuser

Photoreal shots that look directed and sound finished, with dialogue, ambience and deliberate camera moves arriving together in one clip