# Veo Canonical page: https://fuser.studio/models/veo Veo in Fuser runs Veo 3.1, 3.1 Fast, 3 and 3 Fast to generate 4-8 second video with audio at up to 1080p from text, start and end frames or references. ## Specs | Spec | Value | | --- | --- | | Versions | Veo 3.1 Fast, Veo 3.1, Veo 3.1 Lite | | Inputs | Prompt (text), First Frame (image), Last Frame (image), Ingredients (up to 3 images), Negative Prompt (text) | | Output | video | | Duration | 4 seconds, 6 seconds, 8 seconds | | Resolution | 720p, 1080p, 4K | | Aspect ratios | 16:9 Horizontal, 9:16 Vertical | | Audio | Generates audio | | Cost | 165–6,613 credits per run | Photoreal shots that look directed and sound finished, with dialogue, ambience and deliberate camera moves arriving together in one clip Creator: Google DeepMind [Roll the shot](https://app.fuser.studio) ## How Veo works Write the shot the way a director would, add keyframes or reference images if the look must hold, and choose a tier. Picture and sound come back together. ### Write the shot Order your prompt like a director: lens and camera move, subject, action, light, then quoted dialogue and SFX lines, or start from a first-frame image. ### Anchor look and motion Add up to three Ingredients images to keep a character or prop recognizable, or set a last frame to define where the shot lands. ### Pick a tier and render Draft on Fast or Lite, then re-run keepers on the full model with 8-second clips at 1080p or 4K. ## What Veo is good at Sound, continuity and control: spoken lines that sync, characters that survive across shots, keyframe reveals, and three tiers that match the stage of your project. ### Dialogue and sound in one pass Quoted lines come back lip-synced, with room tone and effects that match the scene. Name sounds explicitly with an SFX line so ambience is directed, not guessed. ### Ingredients for recurring characters Attach up to three reference images, covering a character, a prop and a location, and carry them into a new shot. Ingredients run at 8 seconds, so pair them with a steady camera instruction. ### Three tiers, one workflow Use Fast as the default workhorse for iterating prompts and dialogue timing, Lite for the cheapest high-volume drafts, and the full tier for final hero renders you have already proven out. ### First and last frame reveals Set an opening and closing image, then prompt the bridging move, like a slow arc or day turning to night. Choose end frames that are a believable journey within 4 to 8 seconds. ## Made with Veo Short, sound-rich scenes across craft, music, architecture, fashion and the deep sea, each prompt written with a lens, a camera move and an SFX line. ### Craft-commercial frame with directional sound ### Solo performance plate with acoustic resonance ### Architectural light study with room tone ### Vertical lookbook clip with a spoken cue ### Underwater film still with aquatic ambience ## What people build with Veo Filmmakers, commercial directors, fashion teams, social creators and visualization studios reach for it when a clip has to look directed and sound finished. ### Independent filmmakers Block out a spoken scene or an establishing shot with lens, camera move and ambience in the prompt, and get a clip that feels directed. Each clip runs up to 8 seconds. ### Commercial directors Build product hero shots with matched ambience, prove them on Fast, then re-render the keepers on the full tier at 1080p or 4K for client-facing spots. ### Fashion lookbook teams Keep a garment and a model recognizable from clip to clip by feeding Ingredients references into locked 8-second shots. ### Vertical social creators Compose natively in 9:16 with sound already in the clip, so a spoken hook and a room-tone bed arrive without a separate edit. ### Architectural visualization studios Use first and last frames to travel from an exterior render into an interior one, with light and material shifts carried along the camera path. ## Veo questions, answered ### What's the difference between the full, Fast and Lite versions? The full tier is for final hero renders, Fast is the default workhorse for iteration, and Lite is the most affordable option for drafts. The full tier gives the most fidelity but costs the most and renders slowest. Fast gives up some fidelity for speed and a lower cost. Lite costs least but has less polish, so check that settings like top resolutions or Ingredients behave as you need. ### Which Veo tier should I use for drafting versus final delivery? Draft on Fast or Lite, then re-run your keepers on the full model. Fast suits testing prompts, dialogue timing and keyframe setups. Lite suits high-volume previsualization and animatics. Reserve the full tier for high-resolution, client-facing shots you have already proven out, since it is not worth using for exploratory attempts. ### How do I get characters to speak in a clip? Put the spoken line in quotation marks inside your prompt, for example: A woman says, "We have to leave now." Keep each line to about one breath so it fits the clip. Leave Generate Audio on, and add an SFX line for ambience. Keep to one speaker, since lip-sync and attribution drift when several characters talk. ### How do Ingredients work, and are there limits? Ingredients let you attach up to 3 reference images, such as a character, a prop and a location, to carry into a new shot. They only support 8-second output. Use clean, uncluttered references and put the clearest face first. Avoid combining them with a conflicting first frame, which muddies the result. ### How long can a Veo clip be, and what resolutions are available? Clips run 4, 6 or 8 seconds at about 24 fps, in 16:9 or 9:16. Resolution options are 720p, 1080p and 4K, but 1080p, 4K and Ingredients only support 8 seconds. For anything longer than 8 seconds in a single take, this is the wrong tool. ### Can Veo put readable text or subtitles in a video? Not reliably. On-screen text and subtitles tend to come out as garbled lettering, so add captions in post. You can also list terms like subtitles or text overlay in the negative prompt. Crowded multi-person dialogue and strict frame-level motion control are other weak spots. ## Try Veo on Fuser Photoreal shots that look directed and sound finished, with dialogue, ambience and deliberate camera moves arriving together in one clip [Roll the shot](https://app.fuser.studio) ## Related models - [MiniMax H3 Recast](https://fuser.studio/models/minimax-h3-recast.md) - [Kling O3 Edit](https://fuser.studio/models/kling-o3-edit.md) - [Kling O3](https://fuser.studio/models/kling-o3.md) - [Kling 3.0 Video](https://fuser.studio/models/kling-3-0-video.md) - [Seedance 2](https://fuser.studio/models/seedance-2.md) - [MiniMax H3](https://fuser.studio/models/minimax-h3.md) Node reference: [Veo docs](https://docs.fuser.studio/docs/nodes/video/veo.md) ## More articles - [Kling O3 vs Veo 3.1: Same-Prompt Video Test](https://fuser.studio/articles/kling-o3-vs-veo-3-1.md) - [First and Last Frame Video: How to Generate a Clip Between Two Images](https://fuser.studio/articles/first-last-frame-video-guide.md) - [Best AI Avatar Video Generators: One Portrait, One Script, Four Models](https://fuser.studio/articles/best-ai-avatar-video-generators.md)