One click, whole workflow
Recipes are here. Package a whole workflow and run it in one click
Explore recipesOne click, whole workflow
Recipes are here. Package a whole workflow and run it in one click
Explore recipesThree identical tests through LTX-2.5 and Wan 2.6: a potter who speaks one line, a product push-in from a still, and a dialogue shot from a start frame. What each model did, what it costs, and which one you can download.
All guides · Wan 2.6 guide · LTX-2.5 guide
Quick answer: in our three tests, LTX-2.5 was the better pick when a character has to deliver a line to camera: both of its text-to-video runs framed her face for the line and delivered it to the lens. Wan 2.6 was the better pick for a product shot, with a clear push-in, the water drop we asked for and a crisp label, and it followed a start frame's composition more faithfully. Prices are close at 1080p: per second, Wan 2.6 costs a little more than LTX-2.5 Fast and a little less than LTX-2.5 Pro. The bigger difference is ownership: LTX-2.5 has downloadable weights under Lightricks' community licence, while Wan 2.6 is API-only. One run per model per test, so read this as a guide to tendencies, not a ranking.
We reused three Wan 2.6 clips from our Wan 2.6 guide, generated on 28 September 2026, and ran LTX-2.5 on exactly the same prompt text and start frames on 29 September. Each model got one generation per test; nothing was retried or picked from several takes.
Models: Wan 2.6 Video with Multi-Shots off and prompt expansion on, and LTX-2.5 in its Pro mode for all three tests, plus its Fast mode on test 1. These are the versions Fuser's Wan 2.6 Video and LTX 2.5 nodes call.
Settings: 1080p, 16:9, sound on. Durations could not match: Fuser's Wan 2.6 node offers 5, 10 or 15 seconds and LTX-2.5 starts at 6, so Wan ran 5 seconds and LTX-2.5 ran 6. In the side-by-side clips, Wan holds its last frame for the final second. Wan returned 30 fps and LTX-2.5 25 fps.
Test 1, text-to-video with dialogue: "Handheld medium shot of a ceramicist at a pottery wheel shaping a tall vase from wet clay, her hands glistening as the clay spins. She glances up at the camera, smiles and says: 'Almost there.' Warm workshop light. Audio: steady wheel hum, wet clay sounds, quiet studio ambience."
Test 2, product image-to-video: a 1920 × 1080 photo of a SOLA hand wash bottle, with "Slow push-in on the hand wash bottle standing on the stone ledge. Soft sunlight shifts slowly across the tiled wall and a single drop of water runs down the side of the bottle. The label stays sharp and readable. Calm, premium product film. Audio: quiet bathroom ambience and one faint drip."
Test 3, dialogue from a start frame: a 1920 × 1080 night-market still of a woman with a bicycle, with "Handheld shot on a rainy night market street. The woman with the bicycle turns her head to the camera and says: 'I know a better place for noodles.' Rain keeps falling, people with umbrellas walk past behind her, and neon signs reflect on the wet road. Audio: steady rain, distant street chatter, and her voice."
We compared frames by eye, checked for cuts with scene-change detection (none of the seven clips had one), measured each soundtrack's average level and transcribed the speech with Whisper to time the lines.
Wan 2.6 built the warm workshop and the tall vase the prompt asked for, and pushed in slowly until the frame ended on her hands and the vase. She looks up and smiles, and Whisper found "Almost there." at 2.3 to 3.2 seconds, over clearly audible wheel and studio sound.
LTX-2.5 Pro framed her from the waist up with her face in shot for all six seconds. She looks up, talks to the lens and says "Almost there." at 4.0 to 4.9 seconds. The light was cooler than "warm workshop", the vase came out squat rather than tall, and a strange clay-covered post stands beside the wheel for the whole clip.
LTX-2.5 Fast went much closer: she leans her head toward the camera over a medium-height vase and says the line at 4.2 to 5.1 seconds. It had no stray objects, and it costs less than Pro.
Both LTX-2.5 soundtracks averaged about -27 dB against Wan's -17 dB, so expect to raise LTX-2.5's level in the edit.
Verdict for a speaking character: LTX-2.5. Both runs had her face on screen for the line and delivered it to the lens; Fast was the cleaner of the two here. Wan 2.6 was more literal about the set and the vase.
Wan 2.6 gave the push-in we wrote: the bottle grows steadily until the label fills the frame, and a clear drop runs down the side of the bottle. "SOLA" and "HAND WASH" stayed sharp. The start frame's small print is not legible, and Wan redrew it as readable text ("BURGAMOT & SAGE", misspelled), with the smallest lines garbled.
LTX-2.5 Pro moved the camera sideways around the bottle with only a slight push, and no drop appeared. Sunlight and leaf shadows move on the wall, "SOLA" and "HAND WASH" stay sharp, and the small print is garbled.
The LTX API's image-to-video endpoint also takes an optional camera-motion preset such as "dolly_in" (LTX API reference). Fuser's LTX 2.5 node does not expose it, and we did not test it, so in Fuser write the move into the prompt.
Verdict for a product move: Wan 2.6. It carried out both actions in the prompt; LTX-2.5 drifted and skipped the drop.
Wan 2.6 kept the start frame's wide composition: she stands with the bike, turns to the camera and says "I know a better place for noodles." between 0.4 and 2.5 seconds, while shoppers with umbrellas walk past behind her.
LTX-2.5 Pro began on the start frame, then moved in to a close-up of her face; the bicycle drops out of frame. She says the full line straight to the lens between 3.6 and 5.8 seconds, with rain and street sound around it.
Verdict: split. Wan 2.6 followed the prompt and the start frame more literally. LTX-2.5 gave a stronger close-up line reading but reframed the shot. If the start frame's composition has to hold, pick Wan 2.6.
Price at 1080p: per second, LTX-2.5 Fast is the cheapest and LTX-2.5 Pro costs the most (LTX pricing), with Wan 2.6 in between. Because Wan's shortest clip is 5 seconds and LTX-2.5's is 6, our Wan clip was the cheapest of the three, just under the LTX Fast clip, and the LTX Pro clip cost about a third more than either. The order of per-second rates is the same at 720p.
Speed: we did not benchmark generation time. Lightricks describes Fast as built "for speed and low cost" and Pro as "for higher fidelity" (LTX-2.5 model docs); the Fast mode is the one to use for drafts.
Length and resolution in Fuser: the Wan 2.6 node makes 5, 10 or 15 seconds at 720p or 1080p. The LTX 2.5 node makes 6 to 20 seconds in Fast mode, and up to 1440p or 2160p for clips of 10 seconds or less; Pro is limited to 720p or 1080p and 10 seconds.
Framing: Wan 2.6 offers 16:9, 9:16, 1:1, 4:3 and 3:4. LTX 2.5 offers 16:9 and 9:16, or Auto to follow the start image.
Sound: both generate audio with the picture. The LTX 2.5 node's Generate Audio toggle is off by default, so switch it on. Wan 2.6 has no audio switch (every clip we made came with sound) and can also take an audio clip as background music.
Editing: connect a video to the LTX 2.5 node and it switches to retake, which replaces the picture, the sound or both over a time range from a prompt. That mode runs on LTX-2.3's retake endpoint, not LTX-2.5. Wan 2.6 instead offers reference-to-video (up to three clips) and multi-shot prompts with timed cuts; see the Wan 2.6 guide.
LTX-2.5 has downloadable weights. Lightricks publishes the 22B model on Hugging Face (access requires accepting its terms) with code in its LTX-2 GitHub repository, under the LTX-2.x Community License. That licence covers "all LTX-2.5 versions released since August 11, 2026" and requires entities above an annual-revenue threshold set in the licence to obtain a paid licence. Wan 2.6 has no public weights: the Wan-AI Hugging Face organisation has Wan 2.1 and 2.2 repositories but none for Wan 2.5 or 2.6, so Wan 2.6 is used through Alibaba Cloud's API and other hosted APIs. For more models you can run yourself, see the best open-source video models.
In Fuser, connect one prompt node, and a start image if you have one, to a Wan 2.6 Video node and an LTX 2.5 node on the same canvas. Turn on Generate Audio in the LTX node and turn off Multi-Shots in the Wan node for a single take. Draft in LTX-2.5 Fast, then rerun the prompt that works in Pro or send the chosen clip to an upscaler such as Topaz. For more models with sound, see the best AI video generators with audio, or compare three other models in our Seedance, Kling and Veo test.
Three tests at 1080p with sound, one run per model. Wan 2.6 clips generated 28 September 2026 (5 s); LTX-2.5 clips 29 September 2026 (6 s).
| Test | Wan 2.6 | LTX-2.5 |
|---|---|---|
| Results | ||
| Person speaking (text-to-video) | Tall vase and warm set as written; line at 2.3 s; ends on her hands. | Winner. Face in frame and line to camera at about 4 s in Pro and Fast; Pro added a stray object. |
| Product push-in (image-to-video) | Winner. Clear push-in, water drop, sharp headline text. | Sideways drift, slight push, no drop; headline text sharp. |
| Dialogue from a start frame | Kept the wide framing and passers-by; line at 0.4 s. | Pushed in to a close-up line reading at 3.6 s; bicycle left frame. |
| Price per second at 1080p | Between LTX-2.5 Fast and Pro | Fast is the cheapest of the three; Pro costs the most |
| Clip length in Fuser | 5, 10 or 15 s | 6 to 20 s (Fast); up to 10 s (Pro) |
| Open weights | No, API only | Yes, LTX-2.x Community License |
It depends on the shot. In our test LTX-2.5 was better at a character speaking to camera, while Wan 2.6 followed a product camera move and a start frame's composition more literally. Prices per second at 1080p are close to each other.
Its weights are downloadable from Lightricks on Hugging Face under the LTX-2.x Community License. Use is free for smaller entities; companies above the licence's annual-revenue threshold need a paid licence from Lightricks.
No public weights. Alibaba's Wan-AI organisation on Hugging Face has Wan 2.1 and Wan 2.2 repositories but none for Wan 2.6. It is available only as a hosted API, from Alibaba Cloud and other providers, which is how Fuser runs it.
At 1080p, LTX-2.5 Fast is the cheapest per second, Wan 2.6 costs a little more and LTX-2.5 Pro costs the most. Because Wan's shortest clip is 5 seconds and LTX-2.5's is 6, the shortest Wan clip costs slightly less than the shortest LTX Fast clip, and the shortest LTX Pro clip costs about a third more than either.
Yes, both generate sound with the picture, including dialogue. In Fuser, switch on Generate Audio in the LTX 2.5 node, since it is off by default. Wan 2.6 has no audio switch, and every Wan clip we made came with sound.
Fuser's Wan 2.6 node offers 5, 10 or 15 seconds and LTX-2.5's shortest option is 6 seconds, so there was no matching length. We ran Wan at 5 seconds and LTX-2.5 at 6.
One prompt and start frame into Wan 2.6 and LTX-2.5, side by side, then keep the take that fits.