# First and Last Frame Video: How to Generate a Clip Between Two Images

Canonical page: https://fuser.studio/articles/first-last-frame-video-guide

Which AI video models take a start and end frame, how to make a matching pair with an image model and an edit, and what Veo 3.1, Kling 3.0, Seedance 2.5 and MiniMax H3 did with the same pair.

[All guides](https://fuser.studio/articles) · [Best image-to-video models](https://fuser.studio/articles/best-image-to-video-models)

**Quick answer:** to generate a video between a start and an end frame, give an image-to-video model both stills and a prompt that describes the change between them. Among Fuser's current video models, seven take an end frame: Veo 3.1, Kling 3.0, Seedance 2.5 and 2.0, MiniMax H3, Luma Ray 3.2, Vidu Q1 and FLUX.3. Every one of them needs the start frame too; the end frame only counts alongside it. Make the pair by generating the first frame with an image model, then editing that image into the end state, so the camera, subject and set stay identical. In our test, Veo 3.1 Fast, Kling 3.0 and Seedance 2.5 all landed on both frames and played the changes in the order we wrote them; MiniMax H3 hit both frames but lit the candle first, until we spelled the order out.

## Which models take a first and last frame

Every node below needs a start frame for the end frame to count. Most reject an end frame on its own; without a start image, Seedance switches to text-to-video, which has no end-frame input. The input names are the ones on each node in Fuser.

- **[Veo 3.1](https://fuser.studio/models/veo)** (Veo 3.1 and Veo 3.1 Fast): **First Frame** and **Last Frame**. Adding both switches the node to Google's first-and-last-frame mode ([Gemini API: Veo 3.1](https://ai.google.dev/gemini-api/docs/veo)). Only the 3.1 models take a last frame; Veo 3 does not. In Fuser, any Veo run with an image is 8 seconds, at 720p or 1080p.
- **[Kling 3.0 Video](https://fuser.studio/models/kling-3-0-video)** (Standard or Pro): **Image** and **End Image**, 3 to 15 seconds. Kling lists start and end frames among Kling VIDEO 3.0's modes ([Kling VIDEO 3.0 guide](https://kling.ai/quickstart/klingai-video-3-model-user-guide)). Audio is off by default on the node.
- **[Seedance 2](https://fuser.studio/models/seedance-2)** (2.5, and 2.0 Standard, Fast and Mini): **End Image**, used only when you give a single start image. With several images the node switches to reference mode, where the end image doesn't apply. ByteDance's API docs describe the mode as taking "two images as the first and last frames" ([BytePlus ModelArk: Seedance 2.5 tutorial](https://docs.byteplus.com/en/docs/ModelArk/2607688)).
- **[MiniMax H3](https://fuser.studio/models/minimax-h3)** (H3 and H3 Max): **Start Image** and **End Image**, 5 to 15 seconds ([MiniMax API docs: video generation](https://platform.minimax.io/docs/guides/video-generation)). Reference images, video or audio switch the node to reference mode, which doesn't take a start image, so the end frame only works in image-to-video mode.
- **[Luma Ray](https://fuser.studio/models/luma-ray)** (Ray 3.2, generate mode): **End Image**. Luma's API takes an end frame as the clip's last frame; with an end frame the clip is 5 seconds, and it can't be combined with **Loop** ([Luma API: video generation](https://docs.agents.lumalabs.ai/guides/videos/generation)). Our [Luma Ray 3.2 guide](https://fuser.studio/articles/luma-ray-3-2-guide) tests it on a push-in.
- **[Vidu](https://fuser.studio/models/vidu)**: **First Frame** and **Last Frame** on the Q1 model, which routes to Vidu's start-end endpoint. The Q2 model rejects a last frame.
- **FLUX.3** video node: **Start Image** and **End Image**, in full and draft quality, 5 to 20 seconds.

Hailuo 2.3, Wan 2.6 and Grok Imagine Video take a start frame but no end frame in Fuser. So does [LTX 2.5](https://fuser.studio/models/ltx-2-5): Lightricks' LTX-2.5 API accepts a last frame on image-to-video ([LTX docs: LTX-2.5](https://docs.ltx.io/models/ltx-2-5)), but Fuser's node doesn't expose it yet, so it can open on your frame but not land on one ([LTX-2.5 guide](https://fuser.studio/articles/ltx-2-5-guide)). To keep one shot going past a model's length limit, see [how to extend AI video length](https://fuser.studio/articles/extend-ai-video-length).

## Step 1: make a pair of frames that match

The model has to invent every frame in between, so anything that differs between your two stills becomes motion. If the vase moves ten pixels or the window changes shape, you get drift or a morph. The reliable way to avoid that is to make the last frame from the first.

1. **Generate the first frame** with an image model, at the aspect ratio of the video you want. We used Nano Banana 2 (Gemini 3.1 Flash Image) at 16:9 and 2K, and put the candle in the first frame, unlit, so it wouldn't have to appear from nowhere.
2. **Edit it into the end state.** Feed the first frame back into an edit model and change only what should change. Our edit prompt began "Show this exact scene about an hour later, at dusk. Keep the camera position, framing… and every object exactly where they are. Change only the light." Google's pattern for a targeted Gemini edit is the same idea: "change only the [element] to [new element]. Keep everything else in the image exactly the same" ([Gemini API: image generation](https://ai.google.dev/gemini-api/docs/image-generation)). See our [Nano Banana prompt guide](https://fuser.studio/articles/nano-banana-prompt-guide) and [FLUX.1 Kontext editing guide](https://fuser.studio/articles/flux-kontext-editing-guide) for edit prompts.
3. **Compare the two before spending on video.** Flick between them: only the intended change should move.

![The first frame, a terracotta vase on a travertine plinth in late-afternoon sun with an unlit candle, beside the last frame: the same scene at dusk with blue light and the candle lit.](https://statics.fuser.studio/cms/fc041c6f-aad1-4663-8148-04d0e546849f)

_Our test pair. The last frame is an edit of the first, so everything except the light and the flame stays put._

In Fuser this is two Gemini Image nodes: the first generates, the second takes the first's output as its image input and edits it. Wire the first node into the video node's start input and the second into its end input. You can connect the same two images to several video nodes at once, which is how we ran the test below.

## Step 2: write the prompt for the change, not the scene

Both frames already describe what the scene looks like. The prompt's job is what happens between them: the camera, the action and its order. We gave all four models the same prompt:

"Locked-off camera, one continuous shot. Time passes from late afternoon to dusk: the patch of sunlight on the wall slowly slides and fades away, the room cools to a soft evening blue, and the candle beside the vase flickers alight and glows warmly. Audio: quiet room tone and faint distant birdsong."

## Step 3: what four models did with the same pair

We ran Veo 3.1 Fast (1080p, 8 seconds, the length Fuser uses whenever a frame is attached), Kling 3.0 Standard (5 seconds, audio on), Seedance 2.5 (720p, 5 seconds) and MiniMax H3 (768p, 5 seconds) on 28 September 2026, one run each with the same frames and prompt.

![Veo 3.1 Fast, Kling 3.0, Seedance 2.5 and MiniMax H3 clips in a grid, each moving from the sunlit first frame to the dusk last frame as the candle lights.](https://statics.fuser.studio/cms/f53b8d44-c2c9-4a84-8edc-94781b5a6bc4)

_One run per model, played from the same start. Veo runs 8 seconds; the others hold their last frame after about 5._

- **All four hit both frames.** We compared each clip's first and last frame against our inputs: Kling's were closest (about 1.5% average pixel difference), then Seedance (about 3%) and Veo (about 4%, rendered at 1080p). H3 was about 7 to 8%, partly because it returns 1344 x 768, a slightly narrower frame than our input.
- **Veo 3.1 Fast** faded the sun off the wall by about four seconds in, let the room go blue, and lit the candle at about 5.3 seconds, then held the finished dusk scene for the last two and a half seconds.
- **Kling 3.0 Standard** kept the sun patch on the wall longer than Veo and Seedance did, then lit the candle at about 3.7 seconds with a flame that flares sideways before settling, the closest of the four to a wick actually catching.
- **Seedance 2.5** turned the room blue fastest, within about two seconds, and lit the candle at about 3.7 seconds, so the flame comes up in an already dark room.
- **MiniMax H3** lit the candle at about 1.5 seconds, while the late-afternoon sun was still on the wall, then dimmed the room around it.

![Frames at 0 to 100 percent of each clip for Veo 3.1 Fast, Kling 3.0, Seedance 2.5 and two MiniMax H3 runs, showing when the light fades and the candle lights.](https://statics.fuser.studio/cms/3646b500-e656-423c-a582-decd28a3ed81)

_Frames at the same share of each clip's length. The bottom row is H3 rerun with the order spelled out: the candle now waits for the dusk._

H3's response explained its order. The endpoint rewrites prompts and returns the rewrite in a field called expanded_prompt. It pinned our first frame to 0.00 seconds and our last to 5.00 seconds, then described the light shifting while, "Simultaneously, the wick of the candle flickers and ignites." Our prompt listed the changes but never said they came one after another. So we ran H3 once more, same frames and settings, with the order spelled out: "First, the patch of sunlight on the wall slowly slides and fades away as the room cools to a soft evening blue. Only then, in the final second, the candle beside the vase flickers alight and glows warmly." This time the rewrite scheduled the flame "At 00:04.000, as the room reaches its dimmest state", and the clip lit it at about 3.3 seconds, after the room had gone blue. It is the only model we ran twice; the grid above shows the first run.

**Sound, measured by level.** The clips on this page play muted, so we checked each track's loudness over time. Veo, Kling and Seedance each generated a steady background track with a short jump in level at about the moment the flame appears. H3's first track was close to silent (mean level about -58 dB) even though its rewritten prompt described birdsong and the candle's "whoosh"; the second run's was clearly audible (about -47 dB).

## What we'd change next time

- **Put the order in words.** If one change must happen before another, say "first" and "only then". That fixed H3's order in our rerun. If a model rewrites your prompt, as H3 does, read the rewrite before judging the clip.
- **Budget the length for the change.** Veo's 8 seconds gave the dusk scene time to settle before the clip ended; at 5 seconds, Kling and Seedance lit the candle with little more than a second to spare.
- **Keep the pair honest.** The four clips agreed on the start and end because the two frames only differed in light. Bigger differences, such as a moved camera or a new object, ask the model to invent more, and give more room for drift.
- **Use it for joins.** A shared frame between two clips (the last frame of one, the first of the next) is the simplest way to cut between shots without a jump. Our [storyboard-to-animatic guide](https://fuser.studio/articles/ai-storyboard-to-animatic) and [consistent characters guide](https://fuser.studio/articles/consistent-characters-ai) build on that.

## Cost of this test

Per second of output on the day, Seedance 2.5 at 720p cost the most, about three times Veo 3.1 Fast. Veo 3.1 Fast (with audio, at 720p or 1080p) cost a little more than Kling 3.0 Standard with audio, and MiniMax H3 at 768p was the cheapest, at less than half Kling's rate. Per clip, the 5-second Seedance run cost almost twice the 8-second Veo run, Kling's clip cost about half of Veo's, and each H3 run about half of Kling's. For more on each model, see the [Veo 3.1](https://fuser.studio/articles/veo-3-1-prompt-guide), [Kling 3.0](https://fuser.studio/articles/kling-3-prompt-guide), [Seedance 2.5](https://fuser.studio/articles/seedance-2-5-prompt-guide) and [MiniMax H3](https://fuser.studio/articles/minimax-h3-prompt-guide) guides.

## End-frame support in Fuser, and what we saw.

Node inputs from Fuser; test results from one run per model on 28 September 2026.

### Video nodes

| Model | End frame in Fuser | In our test |
| --- | --- | --- |
| Veo 3.1 / 3.1 Fast | Last Frame, with a First Frame; 8 s; 720p or 1080p. | Light, then candle at ~5.3 s; held the end state for the last 2.5 s. |
| Kling 3.0 Standard / Pro | End Image, with an Image; 3 to 15 s. | Closest match to both frames; candle at ~3.7 s. |
| Seedance 2.5 / 2.0 | End Image with a single start image; not in reference mode. | Room went blue first; candle at ~3.7 s. |
| MiniMax H3 / H3 Max | End Image, with a Start Image; not in reference mode. | Candle at ~1.5 s, before the light changed; at ~3.3 s after dusk once the order was spelled out. |
| Luma Ray 3.2 | End Image, with an Image; 5 s; no Loop. | Not in this test; see the Luma Ray 3.2 guide. |
| Vidu Q1 | Last Frame, with a First Frame; Q2 has none. | Not tested. |
| FLUX.3 | End Image, with a Start Image; 5 to 20 s. | Not tested. |

## Questions, answered.

### What is first and last frame video generation?

You give a video model two still images, the first and last frame, plus a prompt, and it generates the motion between them. It's also called keyframe interpolation or start-end frame video.

### Which AI video models support a start and end frame?

Among Fuser's current models: Veo 3.1 and 3.1 Fast, Kling 3.0, Seedance 2.5 and 2.0, MiniMax H3, Luma Ray 3.2, Vidu Q1 and FLUX.3. Hailuo 2.3, Wan 2.6, Grok Imagine Video and LTX-2.5 take a start frame only.

### Can I give only a last frame?

No. Every current Fuser video node that accepts an end frame also needs a start frame. Most reject an end frame on its own; without a start image, Seedance switches to text-to-video, which has no end-frame input.

### How do I make the end frame match the start frame?

Generate the first frame, then edit that image into the end state with an edit model such as Nano Banana, asking it to keep the camera, framing and objects and change only what should change.

### Why did the model do things in a different order than I wrote?

A model can merge a list of changes into one simultaneous transition. MiniMax H3's rewritten prompt did exactly that in our test; when we rewrote ours with 'first' and 'only then', H3 put the changes in order.

### How long should a first-to-last frame clip be?

Long enough for the change to finish. In our test the 8-second Veo clip had time to settle on the end state; the 5-second clips finished the change with about a second to spare.

## Wire two frames into four models.

Make the pair with an image node and an edit, connect it to several video nodes and keep the take that lands.

[See image-to-video models](https://fuser.studio/articles/best-image-to-video-models) · [Explore all guides](https://fuser.studio/articles)

## More articles

- [Runway Aleph 2.0 Guide: Edit a Video With a Prompt or One Frame](https://fuser.studio/articles/runway-aleph-2-guide.md)
- [Best AI Avatar Video Generators: One Portrait, One Script, Four Models](https://fuser.studio/articles/best-ai-avatar-video-generators.md)
- [Best AI Video Generators With Audio: Same Dialogue Prompt, Seven Models](https://fuser.studio/articles/best-ai-video-generators-with-audio.md)
