# Wan 3.0 vs Wan 2.6: Four Same-Prompt Tests Canonical page: https://fuser.studio/articles/wan-3-0-vs-wan-2-6 We ran Wan 3.0 on the exact prompts and start frames we gave Wan 2.6: dialogue, a product push-in and timed multi-shot clips. Verdicts per test, new controls and relative cost. [All guides](https://fuser.studio/articles) · [Wan 3.0 guide](https://fuser.studio/articles/wan-3-0-guide) · [Wan 2.6 guide](https://fuser.studio/articles/wan-2-6-guide) **Quick answer:** in four same-prompt tests at 1080p, Wan 3.0 was the better choice for a character speaking to camera and for a timed multi-shot clip, where it matched the cut times we wrote and kept the vase consistent between shots. Wan 2.6 was as good or better for a product push-in and for holding a start frame's wide composition, and it kept more of a product label readable. In Fuser, Wan 3.0 also adds lengths from 2 to 30 seconds, a 480p draft tier, an end frame, and image, video and audio references. Cost per second is the same at 720p; at 1080p Wan 3.0 costs about a third more. One run per model per test, so read the verdicts as tendencies, not a ranking. ## What changes between the two Fuser nodes Alibaba released the Wan2.6 series on 16 December 2025 with clips "of up to 15 seconds" ([Alibaba announcement](https://www.alibabacloud.com/en/press-room/alibaba-unveils-wan2-6-series-enabling-everyone)). It announced Wan3.0 in public beta on 7 August 2026 with clips up to 30 seconds ([Alibaba announcement](https://www.alibabacloud.com/blog/alibaba-unveils-wan3-0-with-twice-as-long-video-outputs-from-a-richer-variety-of-inputs_603439)) and general availability on 24 August ([Alibaba Cloud blog](https://www.alibabacloud.com/blog/wan3-0-at-general-availability-capabilities-benchmarks-pricing-and-the-workflows-it-changes_603505)), although its API reference, last updated 28 September, still calls the model "currently in preview" ([Wan3.0 API reference](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-api-reference)). Alibaba's own release notes compare Wan3.0 with Wan2.7, a version Fuser doesn't run, so the differences below come from the two Fuser nodes, [Wan 2.6 Video](https://fuser.studio/models/wan-2-6-video) and [Wan 3.0 Video](https://fuser.studio/models/wan-3-0-video), and from our clips. - **Length:** Wan 2.6 makes 5, 10 or 15 seconds. Wan 3.0 takes any whole number of seconds from 2 to 30. - **Resolution:** Wan 2.6 offers 720p or 1080p. Wan 3.0 adds 480p, and Alibaba gives 30 fps as Wan3.0's output rate ([Wan3.0 API reference](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-api-reference)). All eight clips in this test came back at 1920 × 1080 and 30 fps. - **Aspect ratio:** Wan 2.6 has 16:9, 9:16, 1:1, 4:3 and 3:4, with 16:9 as the default. Wan 3.0 has the same five plus Adaptive, its default, which picks a ratio from the inputs. - **Inputs:** Wan 2.6 takes a start image or up to three reference videos, plus an optional audio track used as background music (not with reference videos). Wan 3.0 takes a start image plus an optional end image, or up to ten reference images, five reference videos and five reference audio clips; the videos and the audio clips may each total at most 15 seconds ([Wan3.0 API reference](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-api-reference)). - **Shots:** Wan 2.6 has a Multi-Shots toggle that is on by default. Wan 3.0 has no toggle; Alibaba documents single-shot and multi-shot output from the prompt alone, and says to write "One continuous shot" on the first line when you want no cuts ([Wan3.0 prompt guide](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-prompt-guide)). - **Other controls:** Wan 2.6 has a negative prompt field (up to 500 characters). Wan 3.0 has none; Alibaba's prompt formula ends with a "negative prompt list" written inside the prompt instead ([Wan3.0 prompt guide](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-prompt-guide)). Wan 3.0 adds a Generate Audio switch (on by default), an Enhanced Reasoning toggle (off by default) and a Prime model option. The prompt limit rises from 1,500 characters for Wan 2.6 ([Wan 2.6 API reference](https://www.alibabacloud.com/help/en/model-studio/legacy-wan-text-to-video-api-reference)) to 20,000 ([Wan3.0 API reference](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-api-reference)). ## How we tested We reused the Wan 2.6 clips from our [Wan 2.6 guide](https://fuser.studio/articles/wan-2-6-guide), generated on 28 September 2026, and ran Wan 3.0 on 2 October with exactly the same prompt text and start frames. Both models were run through the same endpoints Fuser's nodes call, and each got one generation per test: nothing was retried or picked from several takes. - **Shared settings:** 1080p, sound on, prompt expansion on. Tests 1 to 3 ran for 5 seconds with Wan 2.6's Multi-Shots off; test 4 ran for 10 seconds with Multi-Shots on. - **Wan 3.0 settings:** the Standard model, Enhanced Reasoning off, 16:9 for the text-to-video tests and Adaptive for the two image-to-video tests (both stills are 1920 × 1080, and both clips came back at that size). - **Checks:** we compared frames by eye, ran scene-change detection to find cuts, measured each soundtrack's average level, and transcribed the speech with Whisper to time the lines. Each clip below opens with a shortened quote of its prompt; the full prompts are in our [Wan 2.6 guide](https://fuser.studio/articles/wan-2-6-guide). The Wan 2.6 clips are the same ones used in our [LTX-2.5 vs Wan 2.6 test](https://fuser.studio/articles/ltx-2-5-vs-wan-2-6). ## Test 1: a person, hands and one spoken line ![Wan 2.6 and Wan 3.0 outputs for the same potter-at-the-wheel dialogue prompt playing side by side, once with each model's sound.](https://statics.fuser.studio/cms/ad1c8796-bf57-49bc-b70f-b3f436aae976) _Test 1, 1080p, 5 s, one run per model. The pair plays twice, first with Wan 2.6’s sound, then with Wan 3.0’s; open it full size to hear it._ - **Wan 2.6** built a warm workshop and a tall, narrow vase, and pushed in slowly until the frame ended on her hands. She looks up and smiles, and Whisper found "Almost there." at 2.3 to 3.2 seconds. - **Wan 3.0** framed her from the waist up in a workshop lined with pots, and her face stays in shot for all five seconds. She looks up into the lens and says "Almost there." at 3.0 to 3.6 seconds, then smiles. The vase came out as a rounded jar rather than tall, and its soundtrack averaged about 6 dB quieter than Wan 2.6's. **Verdict for a speaking character: Wan 3.0.** The line is delivered to camera with her face on screen. Wan 2.6 was more literal about the tall vase. ## Test 2: product push-in from a still ![Wan 2.6 and Wan 3.0 animating the same SOLA hand wash product photo side by side, once with each model's sound.](https://statics.fuser.studio/cms/1afba819-ab49-429c-b4a1-d22a7f6a36a0) _Test 2, same 1920 × 1080 start frame, one run per model. Wan 2.6 generated 28 September, Wan 3.0 on 2 October 2026._ - **Wan 2.6** pushed in until the label nearly filled the frame, and a clear drop runs down the side of the bottle. "SOLA" and "HAND WASH" stayed sharp, and the scent line stayed readable, though misspelled as "BURGAMOT & SAGE". - **Wan 3.0** made the same kind of move, slower and ending wider, and a drop appears on the bottle's shoulder about two seconds in and runs down. "SOLA" and "HAND WASH" stayed sharp; the smaller lines blurred into nonsense letters by the end of the push. Its soundtrack averaged about 9 dB quieter than Wan 2.6's. Small text is a weakness Alibaba itself flags for Wan3.0: "audio texture and on-screen text rendering accuracy are improving but not yet where we want them" ([Alibaba Cloud blog](https://www.alibabacloud.com/blog/wan3-0-30-second-ai-video-generation-from-any-input_603452)). **Verdict for a product move: Wan 2.6, narrowly.** Both carried out the move and the drop; Wan 2.6 pushed in further and kept more of the label legible, and at 1080p it costs less. ## Test 3: dialogue from a start frame ![Wan 2.6 and Wan 3.0 animating the same night market start frame of a woman with a bicycle side by side, once with each model's sound.](https://statics.fuser.studio/cms/f24254e1-aecd-4572-a185-acf52a5a9648) _Test 3, same start frame, one run per model. Speech timings from a Whisper transcript of each clip._ - **Wan 2.6** held the start frame's wide composition: she stands with the bike, turns to the camera and says "I know a better place for noodles." between 0.4 and 2.5 seconds while shoppers with umbrellas pass behind her. - **Wan 3.0** began on the start frame, then moved in to a medium shot with a handheld feel. She turns to the lens and says the full line between 2.2 and 3.4 seconds, then turns away into profile in the last second. The bicycle's handlebars stay in the lower edge of the frame. **Verdict: split.** Wan 2.6 kept the composition you gave it. Wan 3.0 gave a closer line reading and moved the camera the way "handheld shot" suggests, but plan to trim the end of the clip. ## Test 4: three timed shots in one 10-second clip ![Wan 2.6 and Wan 3.0 three-shot pottery clips from the same timed multi-shot prompt, side by side, once with each model sound.](https://statics.fuser.studio/cms/3015496a-ef02-4659-ac1f-c2f90d022995) _Test 4, 10 s at 1080p, one run per model. Cuts: Wan 2.6 at 3.0 s and 5.7 s, Wan 3.0 at 3.0 s and 6.0 s._ Both models got the same prompt: an overall line describing the woman ("early thirties, dark hair tied back, grey linen apron"), then "First shot [0-3s]" wide, "Second shot [3-6s]" a close-up of her hands, and "Third shot [6-10s]" a medium shot where she says "That's the one." Wan 2.6 had its Multi-Shots toggle on; Wan 3.0 has no toggle and got only the prompt. - **Wan 2.6** cut at 3.0 and 5.7 seconds. The woman is the same in the wide and medium shots, but the vase is not: an open, wet grey form in the close-up and a smooth, pale vase in the medium shot. Whisper found the line at 9.0 to 9.7 seconds. - **Wan 3.0** cut at 3.0 and 6.0 seconds. The woman matches across the wide and medium shots, and the tall grey vase in the close-up is recognisably the same one in the medium shot. She says "That's the one." at 9.0 to 9.6 seconds, smiling toward the camera. Alibaba's two Wan3.0 pages disagree on shot length: the usage guide suggests 4 to 6 seconds per shot ([Wan3.0 guide](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-guide)), while the prompt guide says 2 to 5 seconds ([Wan3.0 prompt guide](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-prompt-guide)). Our 3- and 4-second shots worked. **Verdict for multi-shot: Wan 3.0.** It hit the cut times we wrote without a toggle and kept the prop consistent, which Wan 2.6 did not. ## Price - **720p:** the same cost per second on both. - **1080p:** Wan 3.0 costs about a third more per second than Wan 2.6, so each of our 5-second Wan 3.0 clips cost a third more than its Wan 2.6 counterpart. - **480p:** only Wan 3.0 has it, at half its 720p rate, which also makes it half the cost of Wan 2.6's cheapest setting. Alibaba suggests the same pattern of drafting at 480p and finishing at 1080p ([Alibaba Cloud blog](https://www.alibabacloud.com/blog/wan3-0-30-second-ai-video-generation-from-any-input_603452)). - **Prime:** costs 40% more than Standard at 720p and 1080p, and about a third more at 480p. Alibaba describes it as an accelerated version with "significantly faster generation speed" ([wan3.0-video-prime](https://www.alibabacloud.com/help/en/model-studio/wan3-0-video-prime)); we didn't run Prime or time either model. - **Sound:** switching Wan 3.0's audio off does not change the price ([Wan3.0 API reference](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-api-reference)). Cost scales with length on both, so a 30-second Wan 3.0 clip costs six times a 5-second one at the same resolution. ## What we did not test Wan 3.0's headline additions sit outside a like-for-like comparison, because Wan 2.6 has no equivalent: clips longer than 15 seconds, the end frame, and mixed image, video and audio references. Alibaba says the model "natively outputs dialogue, BGM, and sound effects" and supports voice-timbre references ([Wan3.0 guide](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-guide), [Wan3.0 prompt guide](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-prompt-guide)). Alibaba's API also reads documents and web pages as inputs ([Wan3.0 API reference](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-api-reference)); Fuser's Wan 3.0 node does not take files or links. Our [Wan 3.0 guide](https://fuser.studio/articles/wan-3-0-guide) covers the node's inputs on their own. ## Open weights Neither model has public weights. The [Wan-AI organisation on Hugging Face](https://huggingface.co/Wan-AI) and the [Wan-Video organisation on GitHub](https://github.com/Wan-Video) publish Wan 2.1 and Wan 2.2 models (including the Wan2.2-Animate-2 weights) and Wan-Dancer, with no Wan 2.6 or Wan 3.0 repository. Both are used through hosted APIs, which is how Fuser runs them. For models you can download, see [the best open-source video models](https://fuser.studio/articles/best-open-source-video-models). ## Which one to use - **Pick Wan 3.0** for a character who speaks to camera, for multi-shot clips where props must carry over, for anything longer than 15 seconds or shorter than 5, for cheap 480p drafts, and when you need an end frame or reference images, video or audio. - **Keep Wan 2.6** for product shots where small label text matters, when a start frame's composition has to hold, when you want a negative prompt field, and for 1080p finals where its lower per-second cost adds up. To compare on your own brief, connect one prompt node, and a start image if you have one, to a [Wan 2.6 Video](https://fuser.studio/models/wan-2-6-video) node and a [Wan 3.0 Video](https://fuser.studio/models/wan-3-0-video) node on the same canvas. Switch Multi-Shots off in the Wan 2.6 node for a single take, and add "One continuous shot" to the Wan 3.0 prompt if it cuts. Draft Wan 3.0 at 480p, then rerun the prompt that works at 1080p. For a different matchup, see [Wan 3.0 vs Seedance 2.5](https://fuser.studio/articles/wan-3-0-vs-seedance-2-5), or browse [the best AI video generators with audio](https://fuser.studio/articles/best-ai-video-generators-with-audio). ## What we saw, and what changed. Four tests at 1080p with sound, one run per model. Wan 2.6 clips generated 28 September 2026, Wan 3.0 clips 2 October 2026. Specs from Fuser’s two nodes. ### Results | Test or setting | Wan 3.0 | Wan 2.6 | | --- | --- | --- | | Person speaking (text-to-video) | Winner. Face in frame, line to camera at 3.0 s; vase came out squat. | Tall vase as written; push-in ends on her hands; line at 2.3 s. | | Product push-in (image-to-video) | Push-in and drop; headline sharp, small print garbled. | Narrow winner. Closer push-in and drop; more of the label legible. | | Dialogue from a start frame | Pushed in to a medium shot; line at 2.2 s; turns away at the end. | Held the wide framing; line at 0.4 s. | | Three timed shots | Winner. Cuts at 3.0 and 6.0 s with no toggle; same vase across shots. | Cuts at 3.0 and 5.7 s; vase changed between shots. | ### In Fuser | Test or setting | Wan 3.0 | Wan 2.6 | | --- | --- | --- | | Length | 2 to 30 s, any whole second | 5, 10 or 15 s | | Resolution | 480p, 720p, 1080p | 720p, 1080p | | Image inputs | Start and end frame, or up to 10 reference images | Start frame | | Video and audio references | Up to 5 videos and 5 audio clips | Up to 3 videos, or one music track | | Multi-shot | From the prompt, no toggle | Multi-Shots toggle, on by default | | Negative prompt | No field; write it in the prompt | Field, up to 500 characters | | Cost per second | Same at 720p; a third more at 1080p; 480p at half | Lower at 1080p | ## Questions, answered. ### Is Wan 3.0 better than Wan 2.6? In our four tests, Wan 3.0 was better at a character speaking to camera and at a timed multi-shot clip, where it kept a prop consistent between shots. Wan 2.6 was slightly better at a product push-in with readable label text and at holding a start frame’s wide composition. We ran one generation per model per test. ### What does Wan 3.0 add over Wan 2.6 in Fuser? Clips of any whole length from 2 to 30 seconds instead of 5, 10 or 15; a 480p tier; an end frame; up to ten reference images, five reference videos and five reference audio clips; an Adaptive aspect ratio; a Generate Audio switch; an Enhanced Reasoning toggle; and a faster Prime model option. ### Does Wan 3.0 still do multi-shot video? Yes, from the prompt alone. Fuser’s Wan 3.0 node has no Multi-Shots toggle; write each shot with a time range, such as First shot [0-3s]. In our 10-second test it cut at 3.0 and 6.0 seconds, as written. For one continuous take, Alibaba says to start the prompt with One continuous shot. ### Which costs more, Wan 3.0 or Wan 2.6? At 720p they cost the same per second. At 1080p Wan 3.0 costs about a third more per second. Wan 3.0 can also render at 480p for half its 720p rate, cheaper than any Wan 2.6 setting. Wan 3.0 Prime costs 40% more than Standard at 720p and 1080p. ### Is Wan 3.0 open source? No public weights. Alibaba’s Wan-AI organisation on Hugging Face and Wan-Video organisation on GitHub publish Wan 2.1 and Wan 2.2 models but no Wan 2.6 or Wan 3.0 repository. Both models are used through hosted APIs. Alibaba announced Wan3.0 general availability on 24 August 2026, though its API reference still labels the model as in preview. ### Can I use a negative prompt with Wan 3.0? Fuser’s Wan 3.0 node has no separate negative prompt field. Alibaba’s Wan3.0 prompt guide puts a list of things to avoid at the end of the main prompt instead, for example: No subtitles, no watermark. ## Put both Wan models on one canvas. Run one prompt and start frame through Wan 2.6 and Wan 3.0 side by side, then keep the take that fits. [See Wan 3.0 Video](https://fuser.studio/models/wan-3-0-video) · [See Wan 2.6 Video](https://fuser.studio/models/wan-2-6-video) ## More articles - [LTX-2.5 vs Wan 2.6: Same Prompt, Same Start Frame](https://fuser.studio/articles/ltx-2-5-vs-wan-2-6.md) - [Best AI Lip Sync Tools (2026): One Voice Line, Four Models](https://fuser.studio/articles/best-ai-lip-sync-tools.md) - [First and Last Frame Video: How to Generate a Clip Between Two Images](https://fuser.studio/articles/first-last-frame-video-guide.md)