How to Turn a Product Photo into a Video Ad with AI

A step-by-step workflow for turning a single packshot into a short video ad: stage the product, animate the scene, add sound, voiceover and captions, then finish packaging and end card on exact layers.

FuserUpdated
Fuser canvas: a product photo and scene prompt feed a staged product shot, which is animated into a video clip and composed into an end card.

All guides · Chain AI models into one workflow

Quick answer: start from a clean packshot, stage the product in a lifestyle scene with a reference-led image model, animate that still with an image-to-video model such as Seedance 2, Kling 3.0 or Veo, add sound and captions, then put the real packaging and end card on top in the Compositor. Keep every step connected so a new product or a new market only means swapping the input.

A product video ad has two jobs that pull against each other: it needs motion and atmosphere, and it needs the product to look exactly like the thing on the shelf. The workflow below lets AI handle the first and keeps the second under your control.

Step 1: start from a clean packshot

Use a sharp, evenly lit photo of the product on a plain background, shot straight-on. If you only have a busy photo, cut it out first with Background Remover. For our test we generated a fictional product, a frosted hand wash bottle called SOLA, with GPT Image.

Step 2: stage the product in a scene

Connect the packshot to an image model that takes references and describe the set, not the product: "Place this exact bottle, unchanged label and shape, on a pale travertine ledge beside a white basin, fresh sage, morning sunlight and leaf shadows."

Our test run: a GPT Image 2.5 packshot, staged with FLUX.1 Kontext [pro] and animated with Seedance 2.0 Fast. Open it full size to hear the native audio.

Check the product closely before you animate. In our run the staged scene looked right at a glance, but the edit turned frosted glass clear and softened the small label text. That is normal for generative edits, and it is why Step 5 exists.

Step 3: animate the staged still

Image-to-video starts from a frame you already approved, so prompt the motion and ask for the product to stay still: "Slow push-in toward the bottle; sunlight shifts across the wall as leaf shadows sway; the bottle and label stay unchanged."

  1. Seedance 2 uses one image as the first frame, or several as references. Our five-second Seedance 2.0 Fast clip at 720p held the bottle's shape through the push-in and returned its own ambient audio (docs).

  2. Kling 3.0 is strong for cinematic camera moves and can generate audio with the shot.

  3. Veo 3.1 accepts up to three Ingredients, reference images to combine, which helps when the product must appear with a model or a specific location (docs).

Generate in the aspect ratio you will publish: 9:16 for Reels, Stories, Shorts and TikTok, 1:1 or 4:5 for feeds. See social media image sizes for current specs.

Step 4: add sound, voiceover and captions

  • Sound effects: if your clip is silent, connect it to Mirelo SFX or MMAudio, which generate sound synced to the video's motion. Mirelo can take a text prompt to steer it (docs).

  • Voiceover: write the line in a text node and send it to ElevenLabs TTS.

  • Captions: most social video is watched muted. Auto Caption burns subtitles into the video with configurable styling and placement (docs).

Step 5: put the real product back on top

Generative models are not a substitute for your actual packaging. Open the Compositor, place the animated clip as a video layer, and add the original packshot, logo and legal text as their own layers wherever exact detail matters. Finish with an end card: headline, product and call to action. Export MP4 for the ad and PNG for the still variants (export docs). If the clip needs more resolution for large placements, run it through Topaz or SeedVR first.

Make it repeatable

Every step above is a node on one canvas, so the next product is a new packshot on the same graph. Save it as a Recipe and your team gets a product-ad workflow with the packshot and scene as its inputs. For more on campaign stills, see the best AI product photography tools.

The product-ad chain at a glance.

Each step runs as its own node in Fuser, so you can change one without redoing the rest.

StepModels to tryWhat to watch
Image
Clean packshot

Background Remover, GPT Image

Straight-on, even light, plain background.

Stage the scene

FLUX.1 Kontext, Gemini Image, Bria Product Shot

Check label text and materials before animating.

Video and sound
Animate

Seedance 2, Kling 3.0, Veo 3.1

Prompt the motion; ask for the product to stay still.

Sound and voice

Mirelo SFX, MMAudio, ElevenLabs TTS

Add effects to silent clips; write the voiceover as text.

Captions

Auto Caption

Burned-in subtitles for muted autoplay.

Finish

Compositor, Topaz or SeedVR upscale

Overlay real packaging and end card; export MP4.

Questions, answered.

Yes. Stage the product in a scene with a reference-led image model, then animate that still with an image-to-video model such as Seedance 2, Kling 3.0 or Veo. Starting from an approved image keeps the product far more accurate than generating the video from text.

Expect small text and materials to drift in generated edits. Ask the model to keep the product unchanged, check each output, and place your original packshot or label as a separate layer in the Compositor for any shot where exact detail matters.

Use 9:16 for Reels, Stories, Shorts and TikTok, and 1:1 or 4:5 for feed placements. Generate in the ratio you will publish rather than cropping afterwards.

Some video models, such as Seedance 2, can return audio with the clip. For silent clips, connect the video to a video-to-audio model like Mirelo SFX or MMAudio, and use ElevenLabs TTS for a voiceover.

Yes. Save the graph as a Recipe with the packshot and scene description as inputs, and each new product runs through the same staging, animation, sound and finishing steps.

One packshot in. A finished ad out.

Stage, animate, score and finish your product video on one canvas.

All articles