Runway Act-Two Guide: Performance Capture From One Video

How Runway Act-Two works as shipped: the driving performance and character inputs it needs, what transfers with an image versus a video, how to shoot the take, the four settings, and the hard limits.

FuserUpdated
Fuser canvas: GPT Image makes a clay baker; the image and a performance video feed a Runway Act-Two node (1280:720, expression 3, body control on).

All guides · Runway Act-Two in Fuser

Quick answer: Runway Act-Two animates a character from a video of you acting. You give it two things: a driving performance (a 3 to 30 second clip of one person) and a character, as an image or a video. It transfers facial expression to the character, and with a character image it can also transfer hand and body gestures and adds its own camera and background motion. Shoot the performance waist up, well lit, face visible the whole time, with no cuts, and start with your hands in frame. Runway lists these rules in its Act-Two help article; the input limits are in the API reference.

What Act-Two does

Act-Two is performance capture without markers or a rig. It reads the face and body of the person in the driving clip and applies that motion to your character, whether the character is a photo, an illustration, a puppet or a 3D render. Runway says it works with "a range of angles, non-human characters, and styles" (Runway).

It is not a text-to-video model. There is no prompt. Everything about timing, delivery and gesture comes from the take you record, so the quality of the output depends mostly on the quality of that take.

Image or video: what transfers

The character input decides how much of the shot Act-Two controls.

  • Character image. Act-Two drives the face from your performance and, with body control on, your gestures and body movement too. It also adds environmental and camera motion on its own; Runway's example shows a subtle handheld shake. Use an image when you want the performance to carry the whole shot.

  • Character video. The output keeps the subject, background and camera motion from that video. Act-Two replaces the facial movement and expressions, and gesture control is not available. If the character video is shorter than your performance, Runway loops it back and forth ("boomerang") to fill the length, so match the durations when you can.

Pick a video when the camera move or the scene's motion is already right and you only need the face to act. Pick an image for everything else.

Shoot the driving performance

Runway's best practices for the performance clip (Runway):

  1. One person in frame. Act-Two drives a single character per generation.

  2. Waist up at most. Runway asks for the subject framed, at furthest, from the waist up.

  3. Face visible throughout. Don't turn away or cover your face.

  4. Good light and clear expressions. Some expressions are not supported, sticking out your tongue among them.

  5. No cuts. One continuous take.

  6. Hands in frame at the start if you want gestures, ideally with palms facing the camera, and start in a pose close to your character's.

  7. Natural movement. Runway recommends natural movement over excessive or abrupt movement.

Keep the clip between 3 and 30 seconds; the API rejects anything outside that range (Runway API). Runway also asks for clean audio with steady volume and little background noise.

Prepare the character

Two character images built to Runway's checklist with GPT Image 2.5 Flare, the default model in Fuser's GPT Image node. Generated 28 September 2026.

The character needs a recognisable face that stays inside the frame: visible eyes and a mouth, one subject, framed waist up at most (Runway API, Runway). Anything you want to see while it talks has to be visible in the input. Runway's example: if the character should have fangs, use an image where the teeth show.

If you're generating the character rather than drawing it, write those rules straight into the image prompt. Both characters above came from one prompt each in GPT Image: "framed from the waist up, facing the camera directly, both hands raised at chest height with palms facing the camera … eyes clearly visible … single character only". The raised palms mirror the start pose Runway recommends for the performance, so body control has a matching pose to start from. For a character you'll reuse across many shots, see consistent characters.

Runway's API accepts JPEG, PNG and WebP images and MP4, MOV, MKV or WebM video, with input aspect ratios between 0.5 and 2.358 and video files up to 32 MB by URL (Runway input requirements).

Act-Two settings in Fuser

The Runway Act-Two node has three inputs, Character Image, Character Video and Reference Video (the performance), and four controls:

  • Ratio: 1280:720 (default), 720:1280, 960:960, 1104:832, 832:1104 or 1584:672. These are the output sizes in pixels; Runway renders at 24 fps (Runway).

  • Expression Intensity: 1 to 5, default 3. Lower values move the face less but can keep the character more consistent; higher values are more expressive and can introduce artifacts. Runway's advice is to start at 3.

  • Body Control: on by default. When on, gestures and body motion transfer as well as the face. It only has an effect with a character image.

  • Seed: reuse a seed to get a similar result from the same inputs; change it for a different take.

Connect either a character image or a character video, not both; the node stops with an error if both or neither are connected. Outside Fuser, the same model is available directly through Runway's own app and API.

Limits and cost

  • Duration: 3 to 30 seconds, set by the performance clip. The output runs as long as your take.

  • Billing: Runway charges 5 credits per second, so the price scales directly with the length of the take (Runway pricing). Takes shorter than 3 seconds are billed as 3 seconds (15 credits) (Runway). In Fuser, the cost estimate also scales with the length of the performance clip.

  • One character per generation. For a two-person scene, run each character against its own performance and assemble the shots afterwards.

  • Faces first. Runway's input rules ask for a face that stays visible and framing no wider than waist up, so wide shots, full-body action and profiles that hide the face fall outside them.

When to use something else

Act-Two is the right tool when a person can act the scene and you want that acting on a different character. Other jobs have better fits:

Build it as a workflow

In Fuser the character and the performance sit on the same canvas. Generate the character with an image model, drop your recorded take into a Video node, and wire both into Act-Two. From there, send the result to a video upscaler (best AI upscalers) for a sharper master or to VEED subtitles for captions, and keep the character image node in place so the next take uses exactly the same character. The chaining guide covers the pattern in more depth.

Character image or character video?

What Act-Two does with each type of character input.

BehaviourCharacter imageCharacter video
Act-Two behaviour
Facial expression

Transferred from the performance

Transferred from the performance

Gestures and body

Transferred when Body Control is on

Not available; body comes from the character video

Camera and background motion

Added automatically

Kept from the character video

Length mismatch

Not applicable

Character video loops back and forth to match the performance

Best for

Letting the performance drive the whole shot

Keeping an existing camera move or scene motion

Questions, answered.

Act-Two is Runway's performance capture model. It takes a video of a person acting and a character image or video, and transfers the facial expressions, and optionally the gestures and body movement, onto the character.

The driving performance must be between 3 and 30 seconds, and the output matches its length. Runway bills a minimum of 3 seconds.

No. Act-Two has no text prompt. Timing, expression and gesture all come from the performance video, and the look comes from the character input.

Yes, as long as the character has clearly defined eyes and a mouth. Runway says it works with a range of angles, styles and non-human characters.

Gesture transfer needs a character image and Body Control switched on. It is not available with a character video. Start the take with your hands in frame and in a pose similar to the character's.

It sets how much facial motion transfers, from 1 to 5 with a default of 3. Lower values can keep the character more consistent; higher values are more expressive but may add artifacts.

Put your performance on any character.

Generate the character, add your take and run Act-Two on one canvas.

All articles