Wan 3.0 vs Seedance 2.5: Same-Input Video Test

Five tests, one run each: Wan 3.0 and Seedance 2.5 on a product camera move, a spoken line, two start frames and a two-shot prompt. What each produced, which won each test and how their credit costs compare.

FuserUpdated
Fuser canvas: one pottery-wheel prompt wired to a Wan 3.0 Video node and a Seedance 2 node, each showing frames from its real 720p, 6-second output.

All guides · Best image-to-video models

Quick answer: in five same-input tests, Wan 3.0 won three, tied one, and was the only model to return a clip in the remaining test, a start frame of a woman that Seedance 2.5 refused. At 720p it also used about a fifth of the credits. Where Wan 3.0 won, it caught details Seedance 2.5 missed: dust in the light beam, the potter's hands in view while she spoke, and the tall vase the close-up asked for. Seedance 2.5 made the bolder camera push and the more visible water drop. Seedance 2.5's advantages are on paper: three times as many reference images, twice the total length of reference video and audio, and 11 documented speech languages. Each test is one run per model, so read this as each model's tendencies, not a ranking.

How we tested

Tests 1, 2 and 5 reuse Seedance 2.5 clips we generated on 28 September 2026 for our Seedance vs Kling vs Veo test and Seedance 2.5 prompt guide. On 2 October 2026 we ran Wan 3.0 on exactly the same prompts and settings, and ran both models on two new start-frame tests. Every clip is a single generation; nothing was re-rolled or picked from several takes. All clips were made with the same model versions Fuser runs.

  • Models: Wan 3.0 Standard in Fuser's Wan 3.0 Video node, with its defaults of prompt expansion on and enhanced reasoning off, and Seedance 2.5, the default model in Fuser's Seedance 2 node, at standard bitrate.

  • Settings: 720p, audio on, 6 seconds for the text-to-video tests and 5 seconds for the start-frame tests; test 5 ran both at 480p and 5 seconds to match the earlier Seedance clip. Text-to-video was 16:9; start-frame clips take their shape from the frame. Wan 3.0 defaults to 1080p in Fuser, so we set it to 720p to match.

  • Test 1 prompt: "Slow dolly-in on a matte terracotta vase standing on a travertine plinth in a sunlit plaster room. Late-afternoon light slides slowly across the wall as dust drifts through the beam. Calm, premium product film, shallow depth of field. Audio: soft room tone and a faint breeze."

  • Test 2 prompt: "Handheld medium shot of a ceramicist at a pottery wheel shaping a tall vase from wet clay, her hands glistening as the clay spins. She glances up at the camera, smiles and says: 'Almost there.' Warm workshop light. Audio: steady wheel hum, wet clay sounds, quiet studio ambience."

  • Test 3, product start frame: a 1920 × 1080 product image of a SOLA hand wash bottle, with "Slow push-in on the hand wash bottle standing on the stone ledge. Soft sunlight shifts slowly across the tiled wall and a single drop of water runs down the side of the bottle. The label stays sharp and readable. Calm, premium product film. Audio: quiet bathroom ambience and one faint drip."

  • Test 4, a person as the start frame: a 1920 × 1080 photorealistic night-market still of a woman with a bicycle, with "Handheld shot on a rainy night market street. The woman with the bicycle turns her head to the camera and says: ‘I know a better place for noodles.’ Rain keeps falling, people with umbrellas walk past behind her, and neon signs reflect on the wet road. Audio: steady rain, distant street chatter, and her voice."

  • Test 5, two shots in Seedance's format: the storyboard prompt from our Seedance 2.5 guide, written with ByteDance's timestamps and sound brackets (quoted in full below).

We judged frames by eye, found cuts with frame-difference scene detection, measured audio levels and transcribed the spoken lines with Whisper to time them. Wan 3.0 returned 30 fps clips and Seedance 2.5 returned 24 fps, both at 1280 × 720 for the 720p tests.

Test 1: product shot with a camera move

Test 1, one run per model, 720p, 6 seconds. Open it full size to compare the camera move.
  • Wan 3.0 made a slow push-in on a terracotta vase standing on a wide travertine slab, with a band of sunlight moving across a pale plaster wall and visible specks of dust drifting in the beam. One shot, no cuts, and an ambience bed about 11 dB louder on average than Seedance's.

  • Seedance 2.5 made the bigger move: a steady dolly-in from a vase on a block plinth to a tight framing of the vase, and a raking beam of light across a warm wall. Dust was hard to make out, and the room tone was quiet, averaging −42 dB.

Verdict: Wan 3.0, narrowly. It delivered every element of the prompt, including the dust and the moving light. Seedance 2.5's push-in was the more obvious camera move, so pick it if the move is the point of the shot.

Test 2: a person, hands and one spoken line

Test 2, one run per model. Muted here; full size plays Wan 3.0’s sound first, then Seedance 2.5’s.
  • Wan 3.0 held one continuous take with her clay-covered hands on the vase throughout, in a bright, fairly neutral-toned workshop. She looks up into the lens and says "Almost there" at 4.1 seconds, then keeps working and smiling. Her head touches the top edge of the frame when she leans in.

  • Seedance 2.5 also held a single take and delivered the line at about four and a half seconds, looking into the camera, but her hands were mostly hidden behind a tall vase. Her room was darker and warmer in tone, closer to the "warm workshop light" the prompt asked for.

Verdict: Wan 3.0. Both spoke the line in one take; only Wan 3.0 kept the hands, which were half the prompt, in view. Seedance 2.5 came closer on the warm light. Our Wan 3.0 guide covers how Alibaba recommends writing dialogue.

Test 3: product start frame

Test 3, same start frame and prompt, one run per model, 720p, 5 seconds.
  • Wan 3.0 pushed in slowly on the bottle in one shot, with a small drop running down the left edge of the bottle. "SOLA" and "HAND WASH" stayed sharp to the last frame.

  • Seedance 2.5 made a very similar push-in and a larger, more obvious drop running down the right side. "SOLA" and "HAND WASH" also stayed sharp.

  • In both clips the smallest print on the label, which is not legible in the start frame either, came out as garbled lettering.

Verdict: a tie. Both held the frame, the label and the move; Seedance 2.5 showed the drop more clearly. For more on start frames see the first and last frame guide.

Test 4: a person as the start frame

Test 4. Wan 3.0 animated the frame; Seedance 2.5 refused the same request.
  • Wan 3.0 kept the start frame's composition: the woman with her bicycle turns to the camera and says "I know a better place for noodles" between 2.9 and 4.2 seconds, while people with umbrellas walk past and the neon signs stay in place.

  • Seedance 2.5 returned no video. The request was refused with the message "The images or videos provided may contain likenesses of real people or other private information that cannot be processed." ByteDance's documentation states that Seedance 2.5 does not support directly uploading reference images or videos containing real human faces (BytePlus: Seedance 2.5 tutorial). Our product start frame in test 3 went through without a problem.

Verdict: Wan 3.0. If your start frame is a photorealistic person, plan on Wan 3.0; Seedance 2.5 may refuse it. We did not test illustrated or stylised characters.

Test 5: two shots in Seedance's own prompt format

Test 5, 480p, 5 seconds, one run per model. Full size plays both soundtracks in turn.

Prompt: "A ceramicist's studio at dusk, warm tungsten work lamp. Shot 1 [0-2s]: close-up of wet, clay-covered hands shaping a tall vase on a spinning pottery wheel, slow push-in. <steady wheel hum, wet clay sounds>. Hard cut. Shot 2 [2-5s]: medium shot, fixed camera. The ceramicist, a woman in a clay-streaked apron, looks up at the camera and smiles. {English: "Almost there."} (quiet acoustic guitar). Photorealistic, natural colors, shallow depth of field. No subtitles, no text."

This prompt uses ByteDance's conventions: timestamped shot ranges, angle brackets for effects, curly braces for dialogue and round brackets for music (BytePlus: Seedance 2.5 tutorial). Alibaba documents its own format for Wan 3.0, with dialogue written as a speaker, a colon and the quoted line, and timestamped shots that join end to end (Wan 3.0 prompt guide).

  • Wan 3.0 cut at 1.97 seconds, against the [0-2s] boundary, from a close-up of clay-covered hands on a tall cylindrical vase to a medium shot of the potter smiling into the lens under a work lamp, with a dusk window behind her. She says "Almost there" at 3.5 to 4.2 seconds, inside Shot 2. In Shot 2 the pot sits mostly below the frame; the rim that shows matches the cylinder from Shot 1. No subtitles or text appeared.

  • Seedance 2.5 cut at 2.17 seconds and placed the line at 4.3 to 4.9 seconds, also inside Shot 2, but its close-up showed a low, open form rather than the tall vase the prompt asked for, and the pot became a tall cylinder in the medium shot.

Verdict: Wan 3.0. Both hit the timestamps; Wan 3.0 also followed the tall vase in Shot 1 and showed no visible change of object across the cut, on a prompt written for the other model.

What each model offers beyond these tests

  • Release: ByteDance introduced Seedance 2.5 on 31 July 2026 (ByteDance Seed). Alibaba announced Wan 3.0 on 7 August 2026 as a public beta, with testing by application on its Model Studio platform (Alibaba Cloud).

  • Length: both generate up to 30 seconds in one pass. Fuser's Wan 3.0 node takes 2 to 30 seconds and the Seedance node 4 to 30 (Wan 3.0 guide, BytePlus: Seedance 2.5 tutorial).

  • Resolution and frame rate: both offer 480p, 720p and 1080p in Fuser. Alibaba lists 30 fps output for Wan 3.0 (Wan 3.0 guide); ByteDance specifies 10-bit colour depth at 1080p for Seedance 2.5 (BytePlus: Seedance 2.5 tutorial).

  • Aspect ratios in Fuser: Wan 3.0 offers adaptive, 16:9, 4:3, 1:1, 3:4 and 9:16. Seedance adds 21:9. With a start frame, Wan 3.0 follows its shape when the aspect ratio is left on Adaptive, the default; Seedance does so when set to Auto, since its node defaults to 16:9.

  • Start and end frames: both nodes take a start image and an optional end image.

  • References in Fuser: Wan 3.0 takes up to 10 images, 5 videos and 5 audio clips, with video and audio each capped at 15 seconds in total, and references cannot be combined with a start frame. Seedance 2.5 takes up to 30 images, 10 videos and 10 audio clips, with up to 30 seconds of reference audio (ByteDance Seed: Seedance 2.5, BytePlus: Seedance 2.5 tutorial). In Fuser, a Seedance audio reference needs at least one image or video alongside it. In prompts, Wan 3.0 refers to "Image 1" and "Video 1"; Fuser's Seedance node uses @Image1 and @Video1.

  • Speech languages: ByteDance lists 11 languages for Seedance 2.5 prompts and speech: Chinese, English, Spanish, Indonesian, Malay, Thai, Arabic, Portuguese, Vietnamese, Japanese and Korean. Alibaba describes "natural multilingual voice outputs" for Wan 3.0 without publishing a list, and its API reference says prompts support Chinese and English (Alibaba: Wan 3.0 announcement, Wan 3.0 API reference).

  • Extra controls: Fuser's Wan 3.0 node adds Prime, which Alibaba describes as a high-speed version with capabilities aligned to the standard model (Wan 3.0 API reference), plus a prompt-expansion toggle, enhanced reasoning and a seed. The Seedance node adds a high-bitrate option and the older Seedance 2.0, 2.0 Fast and 2.0 Mini models.

  • Weights: we found no downloadable weights for either. ByteDance documents Seedance 2.5 as an API model, and as of 2 October 2026 the Wan-AI organisation on Hugging Face lists Wan 2.1 and 2.2 models but no Wan 3.0 weights (Wan-AI on Hugging Face).

Relative cost

Both nodes charge by output length and resolution, and turning audio off does not lower either price in Fuser. Using Fuser's credit rates at the settings we ran:

  • At 720p: Seedance 2.5 costs about 4.6 times as many credits as Wan 3.0 Standard per second.

  • At 480p: about 4.3 times. At 1080p: about 5.7 times.

  • At each node's default resolution (Wan 3.0 at 1080p, Seedance at 720p): Seedance 2.5 still costs about 2.3 times as much for the same length.

  • Wan 3.0 Prime costs more per second than Standard; we did not test it.

Which one to use

  • Pick Wan 3.0 as the default for text-to-video and start-frame work, especially with people in the frame, for long takes on a budget, or when you want 1080p at a fraction of Seedance's cost. See Wan 3.0 vs Wan 2.6 for what changed from the previous version.

  • Pick Seedance 2.5 when the shot leans on many references (more than 10 images, or more than 15 seconds of reference video or audio), when dialogue needs one of its 11 listed languages, when you need 21:9, or when a bolder camera move matters more than small details, as in our test 1. The Seedance 2.5 prompt guide covers its syntax.

  • Run both on one canvas when in doubt: a single prompt node wired to both video nodes, then send the take you keep to an upscaler or Auto Caption. For more models with native sound, see the best AI video generators with audio.

What we saw, test by test.

One run per model per test, audio on, same prompt and inputs for both.

TestWan 3.0Seedance 2.5
Results
Product dolly-in (text to video)

Narrow winner. Slow push-in, moving light, visible dust, louder ambience.

Bigger, bolder push-in; dust hard to make out; very quiet room tone.

Speaking potter (text to video)

Winner. One take, hands in view, line at 4.1 s; brighter, more neutral light.

One take, line at about 4.5 s, warmer light; hands mostly hidden.

Product start frame

Tie. Push-in, label sharp, small drop.

Tie. Push-in, label sharp, clearer drop.

Person as start frame

Winner. Kept the composition; line at 2.9 to 4.2 s.

Refused: possible real-person likeness.

Two shots, Seedance prompt format

Winner. Cut at 1.97 s, line in Shot 2, tall vase in Shot 1, no visible change across the cut.

Cut at 2.17 s, line in Shot 2; low open pot in Shot 1, tall cylinder in Shot 2.

Relative cost per second at 720p

About a fifth of Seedance 2.5.

About 4.6 times Wan 3.0 (5.7 times at 1080p).

Questions, answered.

In our five same-input tests, Wan 3.0 won three, tied one, and was the only model to animate a start frame showing a person. Seedance 2.5 made the bolder camera move and offers larger reference limits. Each test was one run per model, so treat it as a guide to tendencies.

Wan 3.0. In Fuser, Seedance 2.5 costs about 4.6 times as many credits per second as Wan 3.0 Standard at 720p, about 4.3 times at 480p and about 5.7 times at 1080p.

Both generate up to 30 seconds in one pass. In Fuser, Wan 3.0 takes 2 to 30 seconds and Seedance 2.5 takes 4 to 30 seconds.

ByteDance states that Seedance 2.5 does not support directly uploading reference images or videos containing real human faces. In our test it refused a photorealistic start frame of a woman with a content-policy message, while a product image went through. Wan 3.0 accepted the same frame.

Yes. Both generate audio with the video, including speech, and audio is on by default in both Fuser nodes. ByteDance lists 11 speech languages for Seedance 2.5; Alibaba describes multilingual voice output for Wan 3.0 without publishing a list.

Yes. In Fuser, wire one prompt node, and a start frame if you have one, to a Wan 3.0 Video node and a Seedance 2 node on the same canvas, run both and keep the better take.

Run your brief through both.

One prompt, one start frame, Wan 3.0 and Seedance 2.5 side by side on one canvas.

All articles