One click, whole workflow
Recipes are here. Package a whole workflow and run it in one click
Explore recipesOne click, whole workflow
Recipes are here. Package a whole workflow and run it in one click
Explore recipesWe gave Google Lyria 3.5 and MiniMax Music 3 the same lyrics and the same instrumental brief, twelve tracks in all, and measured what came back: which words were sung, how long each track ran, how loud it was and how it ended.
All guides · Lyria 3.5 · MiniMax Music 3
Quick answer: Both models sang our lyrics, but they handle length and control differently. Lyria 3.5 takes everything in one prompt (style, lyrics, length) and returns an MP3 plus a text map of the song's sections. It costs the same flat amount per song and ran long: asked for 30 seconds it gave about 61, twice. MiniMax Music 3 takes lyrics in a separate field and has Duration, Seed, Steps and Guidance controls. Its Duration works as a ceiling (no run went more than a tenth of a second past it), but it stopped early in four of six runs and was cut off mid-music at the ceiling in the other two. It costs less than Lyria below about 50 seconds and more above. In our runs every Lyria file was louder than every MiniMax file, and five of the six Lyria files ended with about two seconds of silence.
We gave both models the same two briefs, each written in the model's own documented syntax, and generated every output with the same model versions the Fuser nodes run:
Vocal brief. An indie-pop style description (118 BPM, A major, female alto, "the chorus kicks in at 22 seconds", "about 90 seconds long") and 17 lines of our own lyrics: two verses, a chorus sung twice, and an outro. Lyria got the lyrics inside its prompt under a "Lyrics:" header with [Verse 1] and [Chorus] tags, the format in Google's Lyria prompt guide. MiniMax got the same style text as its prompt and the same lines in its Lyrics field with section tags such as [verse] and [chorus] on their own lines, as its model card instructs. MiniMax Duration was 90.
Instrumental brief. A lo-fi hip-hop prompt that says "instrumental only, no vocals", 82 BPM in F major, ending "A 60-second track". We ran it again ending "A 30-second track". MiniMax got the same text, Duration 60 or 30, and the lyrics "[instrumental]", for a reason explained below.
Each brief ran twice per model: 12 tracks in all. MiniMax used the node defaults (30 steps, guidance 1.7) with seeds 7 and 8. Lyria has no seed. Four of the six Lyria tracks came from the identical prompts we ran for our Lyria 3.5 prompt guide; the other two were generated for this comparison. Nobody judged the music by ear. Everything below is measured: duration and format with ffprobe, loudness to the EBU R128 standard, two separate beat trackers for tempo, and a speech-to-text pass to check which of our words were sung.
The table at the end of this page sets them side by side; here is what matters in practice.
Lyria 3.5 is Google DeepMind's music model, announced on 29 July 2026 (Google). In Fuser it has one Prompt input (up to 5,000 characters) and one optional Image Inspiration input. Style, lyrics, structure and length all go in the prompt. There is no duration, seed or tempo setting. It returns a Generated Song (MP3) and a Lyrics text output. Google describes the songs as lasting "a couple of minutes", with duration that "can be influenced through your prompt", and notes that "results may vary between calls, even with the same prompt" (Gemini API docs).
MiniMax Music 3 (MiniMax calls it Music 3.0) was introduced on 13 August 2026 as an open-weights model for complete songs of up to five minutes (MiniMax). Its Fuser node has a Prompt, a separate Lyrics input that is required, and four controls: Duration (1 to 300 seconds, default 60), Inference Steps (1 to 100, default 30), Guidance (0 to 20, default 1.7) and Seed. Duration is described in the node as the "maximum song length"; the model "can finish earlier". It returns a WAV file and the actual duration as a number. MiniMax requires each section tag to sit on its own line, and the Fuser node moves any text written after a tag onto the next line for you.
Both models sang our words. Our speech-to-text pass found, at least in part, 10 of our 12 distinct lyric lines in both Lyria takes, 10 of 12 in MiniMax's seed-7 take and 6 of 12 in its seed-8 take. Speech-to-text misses sung lines, so these are minimums, not scores. We counted the one-line outro, "Fade to gold", only where it was heard on its own, not inside the chorus line that contains the same words. In the seed-8 take it heard almost nothing between 45 and 59 seconds, which is where the second verse would sit; whether that verse was sung unclearly or skipped is for a listener to judge. Lyria's Lyrics output also returned all 17 of our lines, in order, under section labels: [[A0]] for the first verse, [[B1]] for the chorus, A and B again, and [[C4]] for the outro. MiniMax returns no lyrics or section text.
Neither put the chorus at 22 seconds. The first chorus words speech-to-text found were at 32.5 and 25.2 seconds for Lyria, and 16.7 and 30.0 seconds for MiniMax. "The chorus kicks in at 22 seconds" is a timing example from Google's own Lyria prompt guide, yet it was not a reliable control in any of the four takes.
Intros differed. Lyria's vocals entered at around the 8-second mark in both takes. MiniMax's seed-7 take began singing at 0.3 seconds with no intro; seed 8 started at 12.3 seconds.
Length. We asked for "about 90 seconds". Lyria gave 91.2 seconds in one take; in the other the music stopped at 94.1 seconds, then the file held 12 seconds of silence and a quiet tail to 118.7 seconds. MiniMax's seed-7 take ended on its own at 72.3 seconds. Its seed-8 take ran to 90.1 seconds, the Duration ceiling, and its final second was as loud as the seconds before it, so the music was cut off by the ceiling, not finished.
Tempo. We asked for 118 BPM. Both beat trackers agreed on 120 BPM for Lyria's second take and on 122 to 126 BPM for MiniMax's seed-7 take; on the other two takes they disagreed, so we make no claim about those.
Google's Lyria prompt guide says you can "request an instrumental track" in the prompt, and its lo-fi example ends "Instrumental only." MiniMax's hosted API documents an is_instrumental parameter for Music 3.0 (MiniMax API docs), but the MiniMax Music 3 node in Fuser does not expose it, and its Lyrics input cannot be left empty. [instrumental] is one of the section tags MiniMax lists (model card), so we put that tag alone in the Lyrics field, alongside the same "instrumental only, no vocals" prompt Lyria got.
What came back. Speech-to-text found no lyric phrases in any of the four 60-second instrumentals. It returned only stray single words: "you" thirteen times on Lyria's second take, and "you", "so" and "Bye." on MiniMax's two. Speech-to-text cannot tell an instrumental from a wordless vocal line, so listen to the MiniMax result before you use it under dialogue.
Lyria's Lyrics output still returned a section map for the instrumentals, four sections ([[A0]] to [[D3]]) in both 60-second takes, which shows the shape of the track before you play it.
This was the clearest measured difference between the two.
Lyria, "A 60-second track": 66.0 and 60.3 seconds.
Lyria, "A 30-second track": 60.7 and 61.1 seconds. A 30-second request produced a one-minute track both times.
MiniMax, Duration 60: 44.6 and 34.6 seconds, both ending on their own.
MiniMax, Duration 30: 30.0 seconds (cut at the ceiling, last second still at full level) and 19.4 seconds (ended on its own).
So MiniMax never ran more than a tenth of a second past the Duration you set (30.02 and 90.11 seconds were the closest), which suits a fixed slot such as a 15-second ad. But across all six MiniMax runs it either stopped well short (19.4 to 72.3 seconds) or hit the ceiling while still playing. Set Duration a few seconds longer than the slot and plan to trim or fade. Lyria's two 60-second requests landed within six seconds of 60, but a request under a minute is best treated as "about a minute, then cut".
On tempo for the 82 BPM instrumental, where both beat trackers agreed, Lyria measured 83.5 BPM on one take; MiniMax measured 86 to 87 and 89 to 90 BPM on two. On the other five instrumentals the trackers disagreed or both read double time, so we make no claim about them.
Format. Lyria: MP3, 44.1 kHz stereo, about 192 kbps, in all six files. MiniMax: WAV, 44.1 kHz, 16-bit stereo, in all six files. MiniMax's model card says the open-weights model "produces 32 kHz, 16-bit stereo WAV audio" (model card); the version Fuser runs returned 44.1 kHz files every time.
Loudness. Lyria's six files measured -11.0 to -12.6 LUFS integrated; MiniMax's six measured -13.4 to -17.0 LUFS. If you cut between them in one edit, match loudness first; the listening clip above is loudness-matched.
Peaks. Lyria's true peaks ran from -0.1 to +0.5 dBTP, MiniMax's from -0.8 to +0.2 dBTP. Nine of the twelve files measured at or above 0 dBTP, so apply a limiter or normalise before delivery.
Endings. Five of the six Lyria files ended with about two seconds of silence after a fade; the sixth is the take with the 12-second gap described above, which then faded out slowly to the end of the file. No MiniMax file contained a silent stretch of a second or more.
Fuser charges Lyria 3.5 one flat amount per song, whatever its length. MiniMax Music 3 is charged per second of the Duration you set (the ceiling, not the length you get back). The two cost the same at a Duration of about 50 seconds. At 30 seconds MiniMax costs about 60 percent of a Lyria song, at 60 seconds about 1.2 times, at 90 seconds about 1.8 times, and at the 300-second maximum about six times. A long song is cheaper on Lyria; a short cue is cheaper on MiniMax, which is also the one that respects a short length.
Google states that "all generated audio includes a SynthID audio watermark for identification" and that safety filters block prompts requesting "specific artist voices or the generation of copyrighted lyrics" (Gemini API docs). Google's model card says its Generative AI Prohibited Use Policy applies to uses of the model.
MiniMax publishes Music 3's weights on Hugging Face under the MiniMax-Music3 Community License (licence). It requires commercial products or services that use the model to "prominently display 'MiniMax-Music3'" in their user interface, and a separate written authorisation from MiniMax when such products earn more than 20 million US dollars a year. Its acceptable use policy also rules out publishing generated content in public "without clearly and prominently disclosing that such information and/or content is machine-generated". Read the licence against your own use; this is a summary, not legal advice.
A cue that must fit a fixed slot under a minute: MiniMax Music 3. Its Duration setting is a ceiling, and below about 50 seconds it costs less. Generate a couple of seeds and check the ending, since in our runs it either stopped early or was cut at the ceiling.
A full song of two or three minutes: Lyria 3.5. One flat cost however long it runs (we did not test songs that long), and the Lyrics output gives you the words and a section map alongside the audio.
Music from a picture: Lyria 3.5 is the only one of the two with an image input.
An instrumental bed: Lyria 3.5 has a documented way to ask for one. MiniMax's [instrumental] tag gave us no transcribed lyric phrases, but audition it first.
Repeatable variations: MiniMax Music 3 has a Seed control, which the node describes as reusable for reproducible generation. We did not test repeat runs.
Both run as nodes on the same Fuser canvas, so you can generate takes from each next to your image and video nodes and compare them before you commit. For prompting each model in depth, see the Lyria 3.5 prompt guide and the MiniMax Music 3 prompt guide; MiniMax recommends a "Structured Caption" prompt, which this like-for-like test did not use. For writing the lyrics themselves, see how to structure song lyrics for AI; for scoring picture, AI music for video; for how Music 3 compares with MiniMax's earlier model, MiniMax Music 3 vs Music 01; and for the wider field, the best AI music and sound effect generators.
Node controls from Fuser, documented facts from Google and MiniMax, and what we measured on the same briefs (two runs per model per brief).
| Feature | Lyria 3.5 | MiniMax Music 3 |
|---|---|---|
| Inputs and controls in Fuser | ||
| Lyrics | Inside the prompt, under "Lyrics:" | Separate Lyrics input, required |
| Length | Words in the prompt | Duration ceiling, 1 to 300 s |
| Seed, steps, guidance | None | Seed; 1 to 100 steps; guidance 0 to 20 |
| Image input | One image | None |
| Instrumental | Ask for it in the prompt (documented) | Lyrics "[instrumental]" (no node switch) |
| Outputs | MP3 + lyrics with section map | WAV + actual duration |
| Measured on our runs | ||
| Asked 30 s | 60.7 s, 61.1 s | 30.0 s (cut at ceiling), 19.4 s |
| Asked 60 s | 66.0 s, 60.3 s | 44.6 s, 34.6 s |
| Asked about 90 s | 91.2 s; 118.7 s with 12 s of silence | 72.3 s; 90.1 s (cut at ceiling) |
| Lyric lines found by speech-to-text | 10 and 10 of 12 | 10 and 6 of 12 |
| Integrated loudness | -11.0 to -12.6 LUFS | -13.4 to -17.0 LUFS |
| File format | MP3, 44.1 kHz stereo | WAV, 44.1 kHz 16-bit stereo |
| Cost and licence | ||
| Pricing basis | Flat per song | Per second of Duration set |
| Cheaper for | Songs over about 50 s | Cues under about 50 s |
| Watermark / licence | SynthID watermark (Google) | Open weights, MiniMax-Music3 Community License |
It depends on the job. In our runs MiniMax Music 3 stayed within its Duration ceiling and costs less below about 50 seconds, which suits short cues. Lyria 3.5 costs one flat amount per song, returns the lyrics with a section map, accepts an image and was louder, but it ran to about a minute when asked for 30 seconds. Both sang our supplied lyrics.
Yes. Lyria 3.5 takes them inside the prompt under a "Lyrics:" header with tags such as [Verse 1] and [Chorus]. MiniMax Music 3 takes them in a separate Lyrics input with tags such as [verse] and [chorus] on their own lines. Speech-to-text found 10 of our 12 distinct lines in both Lyria takes and 10 and 6 in the two MiniMax takes.
The node has no instrumental switch and requires lyrics. Putting the tag [instrumental] alone in the Lyrics field, with an "instrumental only" prompt, gave two tracks in which speech-to-text found no lyric phrases. Listen before you use them, because speech-to-text cannot rule out wordless vocals.
In Lyria 3.5 you write it into the prompt ("a 60-second track"); our 60-second requests came back at 60.3 and 66.0 seconds, but 30-second requests came back at about 61. In MiniMax Music 3 you set Duration, which is a ceiling: our runs stopped between 19.4 and 72.3 seconds or were cut off at the ceiling while still playing.
Lyria 3.5 is a flat amount per song. MiniMax Music 3 is charged per second of the Duration you set. They cost the same at about 50 seconds; MiniMax is cheaper below that and more expensive above it, up to about six times a Lyria song at the 300-second maximum.
Lyria 3.5 returned MP3 at 44.1 kHz stereo, about 192 kbps. MiniMax Music 3 returned 16-bit stereo WAV at 44.1 kHz in Fuser, although MiniMax's model card lists 32 kHz for the open-weights model.
Put Lyria 3.5 and MiniMax Music 3 side by side on one Fuser canvas, next to the picture you are scoring.