Qwen Image 3 vs Qwen Image 2: Same-Prompt and Edit Tests

Seven identical tests through Qwen Image 3 and Qwen Image 2: a poster, a menu, a five-language sign, a dense newspaper page and three photo edits, with every character counted and a verdict per test.

FuserUpdated
Fuser canvas: one menu prompt wired to Qwen Image 3 and Qwen Image 2, and one café photo with an edit prompt wired to both models, each showing its real output.

All guides · Qwen Image 3 in Fuser · Qwen Image 2 in Fuser

Quick answer: in seven same-input tests, Qwen Image 3 won four, Qwen Image 2 won two and one was a tie. Qwen Image 3 got every character right on a poster and a menu where Qwen Image 2 missed one each, kept the small print sharp when it rewrote a product label, and was the only one of the two to keep the room intact in a three-part photo edit. Qwen Image 2 was exact on a five-language sign, where Qwen Image 3 broke the time, and on the required lines of a dense newspaper page, where Qwen Image 3 added a letter and wrote a story we never asked for. Upgrade for editing and short typography. For long layouts, Qwen Image 3 makes the more convincing page but gives you more text to proofread. At 1K, Qwen Image 3 costs about 14% more than Qwen Image 2 Standard; at 2K it costs about the same as Qwen Image 2 Pro. Each result is one run per model, so read them as tendencies, not a benchmark.

What Alibaba says about Qwen Image 3

Alibaba's Qwen-Image-3.0 announcement, published 22 July 2026, makes four claims that matter for this comparison:

  • Small text. It "supports precise rendering of text as small as 10px".

  • Languages. It "supports native rendering of 12 languages". The post names only three of them in its examples: Japanese, Korean and Spanish.

  • Longer prompts. It "supports up to 4.5k token input", aimed at complex layouts "such as newspapers, storyboards, and exam papers". The Qwen Image 3.0 API reference gives a recommended maximum of 4,500 tokens. For the 2.0 series, Alibaba's Qwen-Image-2.0 announcement mentions 1k-token instructions and its API reference says up to 1,300 tokens.

  • Editing. Both generations generate and edit in one model. The 3.0 API reference accepts one to three input images, with output between 512 × 512 and 2048 × 2048 pixels in total.

Alibaba's API lists two 3.0 models: qwen-image-3.0-pro and qwen-image-3.0, the standard model that "balances quality and speed". Fuser's Qwen Image 3 node has no tier switch, and we found no public statement of which of the two it runs, so we don't claim either. For the 2.0 series, Alibaba describes Pro as having "stronger text rendering, realistic texture, and semantic adherence" than the standard model, and Fuser's Qwen Image 2 node offers both.

What changed in the Fuser node

The two nodes look alike on the canvas. Both switch to edit mode when you connect one to three images, both have a negative prompt (up to 500 characters), an Expand Prompt toggle that is on by default, a Block NSFW toggle, a seed and PNG, JPEG or WEBP output. The differences:

  • Tier vs resolution. Qwen Image 2 has a Model switch: Standard (the default) or Pro. Qwen Image 3 has a Resolution switch instead: 1K (the default) or 2K. At 2K the long edge is 2048 pixels, so a 3:4 image comes back at 1536 × 2048. With the 3:4 preset, both Qwen Image 2 tiers returned 768 × 1024 in our tests.

  • No Auto size on Qwen Image 3. Qwen Image 2 has an Auto size that, in edit mode, keeps roughly the shape of your photo. Qwen Image 3 always uses one of its presets, which come in five shapes (square, plus 4:3 and 16:9 in landscape or portrait), and defaults to square. We set it to Landscape 16:9 for our 1392 × 752 café photo, so it came back at 1024 × 576; Qwen Image 2 on Auto returned 1408 × 768.

  • Longer edit instructions. The Qwen Image 3 node accepts prompts up to 5,000 characters in both modes. On Standard, the Qwen Image 2 edit endpoint Fuser calls caps the instruction at 800 characters.

  • Cost. In Fuser credits, a Qwen Image 3 image at 1K costs about 14% more than Qwen Image 2 Standard, and at 2K about the same as Qwen Image 2 Pro. Pro costs a little over twice as much as Standard.

How we tested

We ran seven tests with identical prompts and input images. Each model ran once per test, with no reruns and no cherry-picking. Both used their Fuser node defaults: Expand Prompt on, Block NSFW on, PNG output. Unless noted, Qwen Image 3 ran at 1K and Qwen Image 2 on Standard, the closest cost match. The Qwen Image 2 poster and menu come from our text-in-images test on 28 September 2026, run at the same 3:4 size and settings. Everything else ran on 2 October 2026, generated with the same model versions Fuser uses. We read every output at full resolution and counted characters that matched the request exactly, spaces excluded.

Tests 1 and 2: a poster and a menu

These are the two typography prompts from our roundup of models for text in images. The poster asks for a two-word headline ("NIGHT BLOOMS"), a subtitle and three lines of small print: a date, a doors-and-tickets line and a lineup that includes the name Theo Lindqvist, 111 characters in all. The menu prompt, verbatim:

"A printed café menu card in flat graphic design: cream paper, dark brown ink, a small coffee cup icon. Title at the top: "CORNER STORE COFFEE". Below it, a price list with exactly these eight items and prices, one per line: "Espresso 3.20", "Cortado 3.80", "Flat White 4.10", "Oat Latte 4.60", "Cardamom Bun 3.90", "Rye Sourdough Toast 5.40", "Pistachio Croissant 4.75", "Iced Hojicha 4.95". At the bottom, in small text: "Open daily 7:30 to 16:00, 14 Wexford Lane"."

Tests 1 and 2, one run per model. Counts are correct characters out of the requested text, spaces excluded.
  • Poster: Qwen Image 3 111/111, Qwen Image 2 110/111. Qwen Image 2 set the hyphen in the tickets line as a dash; Qwen Image 3 kept it as typed. Qwen Image 3 added a small decorative emblem with no text in it.

  • Menu: Qwen Image 3 172/172, Qwen Image 2 171/172. Qwen Image 2 printed the closing time as "16:90". Qwen Image 3 set the prices in a right-aligned column with dot leaders and got the footer right.

Verdict: Qwen Image 3, twice. The margin is one character per prompt, but both misses were the kind a quick look won't catch.

Test 3: five languages on one sign

Alibaba's post names only three of the twelve languages it claims for Qwen Image 3, so we tested those three plus Chinese and English. Prompt:

"A hand-painted wooden welcome sign at the entrance of a small seaside guesthouse, photographed straight on in soft daylight. The sign has five lines of white painted lettering, one per line, in this order: "Welcome", "Bienvenidos", "ようこそ", "환영합니다", "欢迎光临". Under them, a smaller line: "Check-in from 15:00". Nothing else is written on the sign."

Test 3. Both models wrote all five scripts correctly; Qwen Image 3 broke the time.

Both models wrote all five words correctly, in the requested order. The difference was in the small line. Qwen Image 2 got all 48 characters right. Qwen Image 3 got 47: the time reads closer to "15:(0|" than "15:00", with one zero broken into a bracket-like stroke and a stray bar after it. Its sign also sat in a fuller scene, with a shingled cottage and hydrangeas.

Verdict: Qwen Image 2. Qwen Image 2 already handled these five scripts, so Qwen Image 3's language claim didn't show up as an advantage here. We didn't test any other languages.

Test 4: a dense newspaper front page

This is the kind of layout Alibaba says Qwen Image 3 was built for. Our 1,269-character prompt asked for a broadsheet front page with twelve exact lines in quotation marks: a masthead ("The Harbor Ledger"), a dateline with volume and issue numbers, a headline, a subheadline, a photo caption, a second story's headline and first line, a three-line weather box with temperatures, and a section index. That is 352 required characters. The prompt said all other body text could be illegible filler. We ran four settings, as two cost-matched pairs.

Test 4, one run per setting. The 1K Qwen Image 3 run is the cost match for Qwen Image 2 Standard; the 2K run is the cost match for Pro.
  • Qwen Image 2 Standard: 352/352. Every required line exact, the filler left as unreadable texture as allowed.

  • Qwen Image 3 at 1K: every required character present, plus one extra letter: the caption reads "MV Corinine". It split each weather line in two and filled the columns with readable sentences, including invented names, figures and quotes, plus a second headline story ("Harbor Marina Upgrades Approved") that we didn't ask for.

  • Qwen Image 2 Pro: 351/352. The "y" in "Saturdays" came out malformed, closer to "Saturdajs".

  • Qwen Image 3 at 2K: one extra letter again, this time in the subheadline ("approsves"), and again an unrequested story ("Harbor Fest Announced"). At 1536 × 2048 its body copy is legible at full size and reads like a real page, with its own typos.

Every miss in the 352 required characters, cropped from the full-resolution files.

Verdict: Qwen Image 2 for accuracy; Qwen Image 3 for a convincing mock-up. Qwen Image 3 made the more realistic front page at both resolutions, which is what Alibaba's newspaper claim describes. But on the text we actually specified, Qwen Image 2 Standard was exact and Qwen Image 3 was not. Its readable filler also means more invented text to check before anything ships. Pro did not beat Standard here.

Test 5: replace one line on a label

The label test from our Qwen editing guide, on the same product photo: "Change the label text "BERGAMOT & SAGE" to "CEDAR & FIG". Keep the same font, size, colour and position. Change nothing else."

Test 5, same source photo and instruction for both models.

Both models changed the line correctly and left the bottle, pump and background alone. The difference is in the small print. Qwen Image 2 smeared "EFFECTIVE" in the "Gentle · Natural · Effective" line and blurred "10.1 fl oz". Qwen Image 3 kept both lines as sharp as the source, which fits Alibaba's small-text claim.

Verdict: Qwen Image 3.

Test 6: three changes in one instruction

The edit from our image editing roundup, on the same café photo: "Change her yellow rain jacket to a charcoal-grey wool coat. Make it a rainy evening outside the window, with wet street reflections. Write the words OPEN LATE in white painted letters on the window glass. Keep her face, hair, pose, the coffee cup, the table and the bicycle exactly the same."

Test 6. One instruction with three changes and a list of things to keep.

Qwen Image 2 made all three changes but turned the café into a street-side view, moved the bicycle outside and dropped the bag strap across her chest. It did the same thing in our earlier roundup, so this isn't a one-off. Qwen Image 3 kept the brick pillar, the back wall, the empty tables, the strap and the bicycle, and made the coat and rain changes cleanly. Its one flaw: it put the OPEN LATE lettering at the far left of the window, where the frame edge clips the first letters.

Verdict: Qwen Image 3, clearly. It also returned a different shape from the source (see the node section above), so crop or resize before you lay it over the original.

Test 7: swap one object

From the same Qwen editing guide: "Replace the coffee cup on the right with a tall glass of fresh orange juice. Keep the counter, the bread, the croissants, the lighting and the camera angle exactly the same."

Test 7. Both models made the swap and kept the counter.

Both models swapped the cup for a glass of juice in the same spot and kept the loaves, croissants, cloth and window light. Both cropped the small chalkboard sign at the right edge a little tighter than the source.

Verdict: tie. For a single, clearly named swap, Qwen Image 2 Standard is enough.

Which one to use

  • Posters, menus, labels and other short typography: Qwen Image 3 at 1K. It was exact on both prompts and kept fine print sharp in an edit.

  • Edits that must leave the rest of the photo alone: Qwen Image 3. It held the room together where Qwen Image 2 rebuilt it. Pick the Image Size closest to your photo, since there is no Auto option.

  • Simple, single swaps: either. Qwen Image 2 Standard is the cheaper of the two.

  • Long, multi-element layouts: Qwen Image 3 at 2K if you want a page that looks real, then proofread everything. Qwen Image 2 Standard if only the quoted lines matter and you want the filler to stay filler.

  • Times, dates and prices: check them by eye on both. Two of the six character errors across these tests were in a time ("16:90", "15:(0|").

When the final text has to be exact, generate the artwork and set the words as real type in Fuser's Compositor, which has text layers with uploadable fonts and a canvas you set to an exact pixel size. The poster typography workflow walks through that.

Run both on one canvas

Connect one prompt node, or one image plus an instruction, to a Qwen Image 3 node and a Qwen Image 2 node, run them together and compare the results side by side, as in the canvas at the top of this page. For more on writing prompts for the new model, see the Qwen Image 3 guide. To see how it compares with Google's model, read Qwen Image 3 vs Nano Banana.

Tests run on 28 September and 2 October 2026.

Seven tests, one verdict each.

One run per model per test. Qwen Image 3 at 1K and Qwen Image 2 Standard unless noted.

TestWhat happenedWinner
Verdict per test
Jazz poster, 111 characters

Qwen Image 3 111/111. Qwen Image 2 110/111 (hyphen set as a dash).

Qwen Image 3

Café menu, 172 characters

Qwen Image 3 172/172. Qwen Image 2 171/172 ("16:90").

Qwen Image 3

Five-language sign, 48 characters

Qwen Image 2 48/48. Qwen Image 3 47/48, with the time broken.

Qwen Image 2

Newspaper page, 352 characters

Qwen Image 2 Standard exact; Pro one wrong letter. Qwen Image 3 one extra letter at 1K and 2K, plus an unrequested story.

Qwen Image 2 for accuracy

Label text edit

Both changed the line. Qwen Image 2 smeared the small print; Qwen Image 3 kept it sharp.

Qwen Image 3

Three-part photo edit

Qwen Image 2 rebuilt the room and dropped the strap. Qwen Image 3 kept the scene; lettering clipped at the edge.

Qwen Image 3

Single object swap

Both swapped the cup and kept the counter.

Tie

Questions, answered.

In our seven same-input tests, Qwen Image 3 won four, Qwen Image 2 won two and one was a tie. Qwen Image 3 was better at short typography and at edits that should leave the rest of the photo alone. Qwen Image 2 was exact on a five-language sign and on the required lines of a dense newspaper page, where Qwen Image 3 slipped by one character.

Alibaba says Qwen Image 3 natively renders 12 languages and shows Japanese, Korean and Spanish examples. In our test it wrote English, Spanish, Japanese, Korean and Chinese correctly, but Qwen Image 2 did too, and Qwen Image 3 broke the digits in the time on the same sign.

A little. In Fuser credits, Qwen Image 3 at 1K costs about 14% more than Qwen Image 2 Standard. At 2K it costs about the same as Qwen Image 2 Pro.

Pro is the higher tier of Qwen Image 2; Alibaba describes it as having stronger text rendering than Standard. Qwen Image 3 is a newer model; the Fuser node has no tier switch and offers it at 1K or 2K. On our newspaper test, Qwen Image 3 at 2K and Qwen Image 2 Pro each made one character error.

Yes. Connect one to three images to the Qwen Image 3 node and it switches to edit mode. Unlike Qwen Image 2, it has no Auto size and defaults to square, so choose the shape closest to your photo.

Alibaba recommends a maximum of 4,500 tokens. The Fuser node accepts up to 5,000 characters, for generation and editing alike. On Standard, the Qwen Image 2 edit endpoint Fuser uses caps instructions at 800 characters.

Run both Qwen models on one canvas.

Wire one prompt or one photo into Qwen Image 3 and Qwen Image 2 and keep the result that gets your text right.

All articles