# MiniMax H3 Recast Guide: Replace the People in a Video From Photos Canonical page: https://fuser.studio/articles/minimax-h3-recast-guide Replace the people in a 5–30 s video with people from photos using MiniMax H3 Recast: limits, left-to-right mapping, prompts, voice and wardrobe caveats, and a real test. [All guides](https://fuser.studio/articles) · [MiniMax H3 Recast in Fuser](https://fuser.studio/models/minimax-h3-recast) **Quick answer:** MiniMax H3 Recast replaces the people in a video you already have with the people in your photos, and keeps the motion, camera, cuts and sound of the original. Give it a 5 to 30 second clip with no single shot longer than 15 seconds, plus one photo per new person, up to four. With no prompt, photo 1 replaces the leftmost main person, photo 2 the next, and so on, in every shot they appear in. In our test on a two-shot, 5-second clip, both swaps landed in the right places on the first run, the cut stayed on the same frame and the soundtrack came back unchanged. Two things to plan for: the new people wore the clothes from their photos, and the original voices stay, so a line spoken by one person is now spoken, in their voice, by whoever replaced them. ## What H3 Max Recast is Recast runs on H3 Max, a post-trained variant of MiniMax H3. MiniMax launched H3 on 31 July 2026 as a general-purpose multimodal generation model that makes video with native stereo sound, up to 15 seconds at 2K, and lists native multi-shot modeling in its training ([MiniMax](https://www.minimax.io/blog/minimax-h3)). MiniMax's own H3 pages describe editing a source video with reference subjects in general terms ([MiniMax](https://www.minimax.io/news/minimax-h3-open-source)), but they do not cover H3 Max or Recast, so every limit below comes from the model's input schema, Fuser's node and our own run, not from MiniMax. In Fuser it is a separate node, **MiniMax H3 Recast**, added on 2 October 2026. It does one job: people in, people out. For generating new shots with H3 and H3 Max, see the [MiniMax H3 prompt guide](https://fuser.studio/articles/minimax-h3-prompt-guide). ## What goes in - **Video** (required): 5 to 30 seconds, and no single shot may run longer than 15 seconds. Its sound is kept. - **People** (required): 1 to 4 photos, one per new person. - **Prompt** (optional): up to 2,000 characters, to say who becomes whom or what else to keep or change. - **Resolution**: 1080p (the default) or 768p. The output keeps the length of the source. The model behind the node accepts a seed, but Fuser's node does not expose one, so two runs with the same inputs can differ. ## How the default mapping works Without a prompt, the photos are matched to the main people in the video from left to right: photo 1 to the leftmost, photo 2 to the next. Each new person replaces their match in every shot that person appears in, so a two-shot scene does not need two runs. That order is the whole contract, so connect the photos in the order the people stand in the frame. We tested that order against a pairing a model might prefer if it matched by look. Our source shows a young man on the left and a woman in her forties on the right. Photo 1 was a woman in her late twenties with braids; photo 2 was a man in his sixties with a white beard and glasses. If the model paired by gender or age it would cross them over. It didn't. ## Our test The source clip is a bakery scene we generated with H3 Max at 768p: two bakers in green aprons behind a counter, a wide shot where he sets down a tray of croissants, then a cut at 2.17 seconds to a closer shot where she turns to the camera and says "Fresh out of the oven." We trimmed it from 5.18 to 5.00 seconds. The two people photos were generated with [FLUX.3 Image](https://fuser.studio/models/flux-3-image). Everyone in this article is AI-generated; no real person appears. We ran Recast once, at 1080p, with no prompt, on 2 October 2026. It did not need a second run. ![A two-shot bakery clip beside its MiniMax H3 Recast output, where a woman with braids and an older bearded man replace the two bakers; reference photos below.](https://statics.fuser.studio/cms/891960d6-eb04-488c-a860-51283d55911e) _MiniMax H3 Recast, 1080p, no prompt, one run on a 5-second, two-shot source, 2 October 2026. Source made with MiniMax H3 Max, photos with FLUX.3 Image; all people are AI-generated. The sound is the recast output’s, which is the source soundtrack unchanged._ ## What it kept, and what it changed **The mapping held in both shots.** The braided woman took the young man's place on the left and the bearded man took the woman's place on the right, in the wide shot and after the cut. Faces, hair and the man's glasses match the photos and stay consistent across the cut. **The edit structure held.** The output is 1920 × 1080 at 24 fps and 5.00 seconds long, against a 1344 × 768, 24 fps, 5.00-second source. Scene detection puts the cut at 2.167 seconds in both files. The framing, the camera, the tray being set down, the chalkboard, the lamps and the window all line up with the source. ![The same frame from the source and the recast: the bakers in green aprons are replaced by a woman in a yellow sweater and a bearded man in a denim shirt, with the counter, tray and room unchanged.](https://statics.fuser.studio/cms/ab7d88e8-68c8-47b5-9421-c3180f4a7945) _The frame at 1.0 seconds from the source (left) and the recast (right). Same run as above._ **The clothes came from the photos.** Both bakers lost their green aprons and now wear the yellow sweater and the denim shirt from the reference photos, with new trousers below the counter line. If wardrobe matters, dress the person in the photo the way they should appear in the scene. The node's help says the prompt can also say what else to keep or change, but we did not test a wardrobe instruction. **Small props can drift.** The rolls on the board in front of the right-hand person changed shape. We did not spot any other change in the room. ## The voice stays with the soundtrack Recast does not touch the audio. The output's soundtrack matched the source's: identical loudness (mean −28.6 dB, peak −10.9 dB) and a waveform correlation above 0.999 with no offset. That includes the line "Fresh out of the oven", which a transcript confirms, and a short greeting in the first shot that H3 Max added on its own when we made the source. The consequence is easy to miss. In the source, the woman speaks the line in the second shot. After the recast, the bearded man in her place mouths the line in her voice. His mouth moves where hers did, so the timing lines up; the voice is simply someone else's. Three ways to handle it: 1. **Map speakers to similar voices.** If the new person can plausibly have the original speaker's voice, nothing else is needed. 2. **Recast non-speaking roles only.** Background people and silent shots carry no voice problem. 3. **Replace the dialogue.** Generate a new line with the ElevenLabs TTS node and re-sync the mouth with a lip-sync node such as [LatentSync](https://fuser.studio/models/latentsync); see [the best AI lip-sync tools](https://fuser.studio/articles/best-ai-lip-sync-tools). We did not test that chain on a recast clip. ## Prompting who becomes whom Use the prompt when the default order is wrong: the people you want to replace are not the leftmost main people, or you want photo 2 on the left. Describe each person in the video by position and something visible, and each photo by what is in it. For our [replace-people test](https://fuser.studio/articles/replace-people-in-video-ai), the prompt that reversed the default order read: "The woman in the yellow raincoat from the second photo replaces the woman with curly hair on the left. The grey-bearded man in the denim jacket from the first photo replaces the young man in the grey hoodie on the right. Keep the coffee cart, the street, the camera movement and every action exactly as they are." The output followed it, and the cart, street and camera stayed as they were. End the prompt with what must not change; the same test gave [Kling O3 Edit](https://fuser.studio/articles/kling-o3-edit-guide) the same closing line. ## Choosing the photos - **One person per photo.** The node takes one photo per new person, not several angles of one. - **Face clear, even light.** We used waist-up portraits on a plain light grey background, and both came through cleanly. We did not test busier photos. - **Dress them for the scene.** In our run the clothing in the photo replaced the costume in the video. - **Large photos are fine.** Fuser scales photos stored in your project down to fit within 1024 × 1024 before sending them; it does not crop or upscale. - **Only people you have permission to use.** See below. ## Preparing the source clip The 5-second minimum and the 15-second shot limit apply to the clip you send, so trim or split long takes first. A 40-second scene has to go through in pieces, each 30 seconds or less. Recast is charged per second of video, and the output always matches the source length. Fuser's estimate rounds the source up to whole seconds: our clip came out of H3 Max at 5.18 seconds, and trimming it to 5.00 seconds brought the estimate from six seconds down to five. ## 768p or 1080p 1080p is the default, and 768p costs two-thirds as much per second. At 1080p, Recast costs more than three times as much per second as a Kling O3 Edit Standard edit, so it is worth getting the photos and the mapping right before running a long clip. Our 768p source, recast at 1080p, came back at 1920 × 1080. ## Recast or another editor Recast is built for one job, replacing whole people while keeping everything else. [Kling O3 Edit](https://fuser.studio/models/kling-o3-edit) can also swap a person from a reference image, as part of a general text-driven editor; on the same two-person clip in our [replace-people test](https://fuser.studio/articles/replace-people-in-video-ai), the two were too close to call on one run each. If you want a new character to perform the motion from a video rather than step into an existing scene, [Kling 3.0 Motion Control](https://fuser.studio/articles/kling-3-motion-control-guide) is the closer fit. The [best AI video editing models](https://fuser.studio/articles/best-ai-video-editing-models) compares the wider field on one source clip. ## Consent and disclosure Recast makes a realistic video of a person doing something they never did. Treat it that way: - **Get permission from the people in the photos** before putting their face into a video, and tell them where it will be used. - **Respect the people you replace.** Their performance, timing and voice stay in the result, so their agreement matters too. - **Never use photos of public figures, colleagues or strangers** without their consent, and never to put words in someone's mouth. - **Disclose the change.** Say that the people were replaced with AI wherever the video is published, and check the labelling rules of the platform and country you publish in. ## Build it in Fuser The canvas at the top is this test: the H3 Max source clip and two FLUX.3 Image portraits feeding a MiniMax H3 Recast node. 1. Add your clip with a Video node, or generate one with [MiniMax H3](https://fuser.studio/models/minimax-h3) or another video node, and connect it to **Video**. 2. Connect one photo per new person to **People**, in left-to-right order. 3. Leave **Prompt** empty to use the default order, or write who becomes whom and what to keep. 4. Choose **Resolution**, then run. 5. To try another cast on the same scene, duplicate the node and swap the photos. To compare H3 against Kling 3.0 on generation rather than editing, see [MiniMax H3 vs Kling 3.0](https://fuser.studio/articles/minimax-h3-vs-kling-3). ## MiniMax H3 Recast at a glance. Inputs and settings on the Fuser node, with what we saw in our run. ### Inputs | Setting | Options | Notes | | --- | --- | --- | | Video | 5 to 30 s, no single shot over 15 s | Required. The sound is kept as it is | | People | 1 to 4 photos, one per person | Required. Default: photo 1 replaces the leftmost main person | | Prompt | Up to 2,000 characters | Optional. Who becomes whom, and what to keep or change | ### Controls | Setting | Options | Notes | | --- | --- | --- | | Resolution | 1080p (default) or 768p | 768p costs two-thirds as much per second | | Length | Same as the source | Charged per second of source, so trim first | | Seed | Not exposed in Fuser | Reruns with the same inputs can differ | ### In our test | Setting | Options | Notes | | --- | --- | --- | | Mapping | Held in both shots | Left to right, despite a cross-gender, cross-age pairing | | Cuts and camera | Kept | Cut at 2.167 s in source and output | | Clothing | Taken from the photos | Dress the photo for the scene | | Voice | Unchanged | New people speak with the original voices | ## Questions, answered. ### What is MiniMax H3 Recast? A video-to-video model, built on MiniMax H3, that replaces the people in an existing clip with the people in reference photos while keeping the motion, camera work, cuts and sound. In Fuser it is the MiniMax H3 Recast node. ### How long can the video be? Between 5 and 30 seconds, and no single shot can be longer than 15 seconds. The recast video has the same length as the source. ### How many people can it replace? Up to four, with one photo per new person. With no prompt they replace the main people in the video from left to right. ### How do I choose who replaces whom? Connect the photos in the order the people stand, left to right, or write a prompt that names each person in the video by position and appearance and says which photo replaces them. ### Does Recast change the voices? No. The soundtrack is kept exactly as it is, so a new person speaks with the original speaker’s voice. Map speakers to similar voices, or replace the dialogue and re-sync the lips afterwards. ### Does it keep the original clothes? Not in our test. Both new people wore the clothes from their reference photos, and the costumes in the source were gone. Dress the person in the photo for the scene. ### Can I recast a real person? Only with that person’s permission, and you should disclose that the video was altered with AI. Never use someone’s photo to make them appear in a video without their consent. ## Recast your scene. Connect a clip and a photo per person to MiniMax H3 Recast and compare casts side by side on one canvas. [Open MiniMax H3 Recast in Fuser](https://fuser.studio/models/minimax-h3-recast) · [Read how to replace people in a video](https://fuser.studio/articles/replace-people-in-video-ai) ## More articles - [How to Replace People in a Video With AI: Recast, Kling O3 Edit and Motion Transfer](https://fuser.studio/articles/replace-people-in-video-ai.md) - [Best AI Video Editing Models: Same Clip, Same Edits](https://fuser.studio/articles/best-ai-video-editing-models.md) - [Kling O3 Edit Guide: Edit a Video With Text and Reference Images](https://fuser.studio/articles/kling-o3-edit-guide.md)