One click, whole workflow
Recipes are here. Package a whole workflow and run it in one click
Explore recipesOne click, whole workflow
Recipes are here. Package a whole workflow and run it in one click
Explore recipesMusic-01 copies a reference song; Music 3 follows a written style prompt and section tags. We gave both the same lyrics and measured length, sung lines, format and loudness.
All guides · MiniMax Music 3 in Fuser · MiniMax Music-01 in Fuser
Quick answer: MiniMax Music 3 is not a drop-in upgrade of Music-01; the two take different inputs. Music-01 (the Minimax Music node in Fuser) copies the feel of a reference song you supply and sings up to 600 characters of lyrics over it, with no style text and no length setting. Music 3 takes a written style prompt, lyrics with section tags such as [verse] and [chorus], and a Duration cap of up to 300 seconds. Given the same lyrics, Music-01 returned 31.5 to 57.1 seconds of audio with every line found in the transcript; with its Duration cap at 60 seconds or more, Music 3 returned 51.4 to 61.1 seconds, with longer stretches where no lyric was transcribed; in one run the first lyric line did not start until 27 seconds in. Choose Music 3 for full-length songs you describe in words; choose Music-01 when you already have a track whose sound you want to reuse.
MiniMax announced Music-01 on 31 August 2024. You "upload a piece of reference music", the model "will automatically learn the rhythm and style of the vocals and accompaniment", and after you enter lyrics you get a new piece of music, "up to 60 seconds long" at launch (MiniMax, Music-01). Fuser's Minimax Music node exposes exactly that: a lyrics field of up to 600 characters, where a new line starts a sung line, a blank line adds a pause and ## at the start and end adds accompaniment, plus a required reference song (.wav or .mp3, over 15 seconds, with music and vocals). There is no seed, no style text and no duration control. Our MiniMax Music prompt guide covers those controls in detail.
MiniMax introduced Music 3.0 on 13 August 2026 as a model that, "given a creative concept and optional lyrics", "composes, arranges, performs, and produces a complete song in a single generation", with "a complete song of up to five minutes" as the target and section tags "such as [intro], [verse], [pre-chorus], [chorus], [bridge], [instrumental], [solo], and [outro]" defining the song's structure (MiniMax, Music 3.0). The weights are public on Hugging Face (model card, which also lists [post-chorus]) under the MiniMax-Music3 Community License, which asks commercial products to display "MiniMax-Music3" in their interface and to get written authorisation above 20 million US dollars of yearly revenue (licence).
Fuser's MiniMax Music 3 node has six controls: Prompt (the style description), Lyrics, Duration (1 to 300 seconds, default 60), Inference Steps (1 to 100, default 30), Guidance (0 to 20, default 1.7) and Seed. Two details are Fuser-specific. The node requires lyrics, although MiniMax describes them as optional. And because the provider discards lyric text written on the same line as a leading section tag, the node moves that text onto its own line before sending the request, so the words in "[chorus] Carry me home" stay in the lyric.
MiniMax's current music API reference lists music-3.0, music-2.6 and music-cover (plus free variants of each), and does not list music-01 (MiniMax API docs). Fuser still runs it, and does not run Music 2.x, so these two are the MiniMax music models you can wire up on a Fuser canvas.
Music-01 needs a reference song, so we made one we own. Music 3 generated a 40-second track from a different lyric ("Morning trains and coffee steam…") and this style prompt: "Genre: indie folk-pop. BPM: 92. Key: G major. Warm and hopeful, building into the chorus. Vocals: clear male lead, close and natural, with light harmonies in the chorus. Arrangement: strummed acoustic guitar and piano; bass and brushed drums enter in the chorus." That track became Music-01's reference, and the same style prompt went to every Music 3 run, so both models were aimed at the same sound.
We then gave both models the same lyric text, formatted for each: an 8-line verse and chorus (251 characters in Music-01 format), and a 16-line version with a second verse and a repeated chorus (498 characters). Music-01 got the lines wrapped in ## with a blank line between sections; Music 3 got [verse] and [chorus] tags. Music 3 ran at Fuser's defaults (Guidance 1.7, 30 steps) with seed 7, plus seed 8 for a second 8-line run, and Duration caps of 60, 120 and 30 seconds as noted. Music-01 ran twice on the 8-line lyric and once on the 16-line lyric. That is eight generations in all, including the reference; every one is reported below, none was re-run or dropped.
We measured each file's format and length, its loudness to the EBU R128 standard, and where each lyric line landed using a Whisper speech-to-text pass. Nobody on our side critiqued the songs by ear, so this page reports measurements, not opinions on vocal or mix quality.
8-line lyric, Music-01: 32.5 and 31.5 seconds. Singing started at about 3 seconds and ended 2.5 to 3 seconds before the end of the file.
8-line lyric, Music 3, 60-second cap: 51.4 seconds with seed 7, which stopped on its own, and 60.1 seconds with seed 8, which ran to the cap. Seed 7 left a gap of more than five seconds between verse and chorus. Seed 8 opened with two wordless "ooh" phrases (at 11.7 and 19.7 seconds) that are not in the lyric and did not sing the first word until 27.2 seconds.
16-line lyric, Music-01: 57.1 seconds, close to the 60-second limit MiniMax stated at launch.
16-line lyric, Music 3, 120-second cap: 61.1 seconds. It stopped well short of the cap. The transcript found no words for about nine seconds between the two halves of the song, and none in the last nine seconds of the file.
The pattern across these runs: Music-01's length tracks the number of lines, at roughly 3.5 to 4 seconds of audio per line. Music 3 produced longer files from the same words, with more time where no lyric was transcribed, and two different seeds gave very different openings. If you need a track to hit a length, Music 3 gives you a cap to set; with Music-01 you change the lyric.
Speech-to-text on sung vocals is imperfect: it misses words, merges lines and mishears ("Vapor lanterns" for "Paper lanterns" in one Music-01 run). So treat these counts as "lines the transcript found", not a final verdict.
Music-01: 8 of 8, 8 of 8 and 16 of 16.
Music 3, 8-line lyric: 8 of 8 in both runs; in the seed 7 run the first chorus line was only partly transcribed.
Music 3, 16-line lyric, 120-second cap: 14 of 16 at least in part. The two lines not found were both "Hold the light and don't let go", the last line of each chorus. Three more were partial: the second "Carry me home, carry me home" came through as one "Carry me home", "Old men fishing in the shallows" as the single word "shadows", and "turns to gold" in the second chorus as "turns to go".
We ran the 16-line lyric through Music 3 again with the same seed and the cap at 30 seconds. The file ran 30.0 seconds. The transcript found the first six lines, placed within about 0.6 seconds of where the 120-second run put them, then a single word, "shallow", half a second before the file ended; line nine is the one that ends in "shallows". The cap cut the song off; it did not fit the lyric into the time. Set Duration above what the lyric needs. Given room, Music 3 took 51 to 60 seconds for our 8 lines and 61 seconds for 16.
Music 3: 16-bit stereo WAV at 44.1 kHz in all five runs. The model card for the open-weights release states 32 kHz output (MiniMax-Music3); the version we ran returned 44.1 kHz files.
Music-01: stereo MP3 at 44.1 kHz, 256 kbps in two runs and 128 kbps in one. The two 8-line runs had identical inputs but came back at different bitrates, and the node has no format setting.
Loudness: Music-01 measured −11.4 to −13.2 LUFS integrated, with a loudness range of 3.2 to 6.3 LU. Music 3 measured −13.4 to −15.2 LUFS, with a loudness range of 5.0 to 14.0 LU. Music-01 came out louder on both lyrics; the loudness ranges overlapped, though Music 3's widest, at 14.0 LU, was more than double Music-01's widest.
Peaks: every file from both models peaked between −0.5 and +0.2 dB true peak, so add a limiter or turn the track down before mixing it under dialogue.
In Fuser, Music-01 is priced per song and Music 3 per second of the Duration you set; the node's own help text notes that the estimate uses this upper bound even if the song ends earlier. At the default 60-second cap a Music 3 run costs about 3.4 times a Music-01 song, and at 30 seconds about 1.7 times. Because Music 3 is priced on the cap, set it close to, but above, what your lyric needs.
Choose Music 3 when you want to describe the sound in words (genre, tempo, key, voice, arrangement), need a song longer than a minute, want section structure from tags, or want a seed you can reuse. Its WAV output is also easier to edit without another lossy encode.
Choose Music-01 when you already have a track with the sound you want, such as your own demo or an earlier generation, and need a new lyric sung in that style in under a minute, at a flat per-song price.
Use both: generate a Music 3 track and feed it to Music-01 as the reference, which is how this test was built. On the canvas that is one wire, from the Music 3 node's Generated Song output to the Minimax Music node's Reference Audio input.
Both nodes output audio that plugs into the rest of a workflow. Wan 2.6 Video accepts an audio input as background music and keeps the first 5, 10 or 15 seconds to match the clip, LatentSync lip-syncs a video to an audio track, and the Whisper node transcribes a song for captions. For writing the prompt and tags, see the MiniMax Music 3 prompt guide and the song lyrics structure guide; for scoring footage, AI music for video; and for a comparison with another music model Fuser runs, Lyria 3.5 vs MiniMax Music 3. The wider pattern is in chaining AI models in one workflow.
What each model takes, and what it returned on the same lyrics.
| Feature | MiniMax Music 3 | MiniMax Music-01 |
|---|---|---|
| Inputs | ||
| Style control | Written style prompt | A reference song (.wav or .mp3, over 15 s) |
| Lyrics | Section tags such as [verse], [chorus] | Up to 600 characters; line breaks, blank lines, ## |
| Length control | Duration cap, 1 to 300 s | None; follows the lyric |
| Other settings | Seed, guidance, inference steps | None |
| Measured in our runs | ||
| 8-line lyric | 51.4 s and 60.1 s (60 s cap) | 32.5 s and 31.5 s |
| 16-line lyric | 61.1 s (120 s cap) | 57.1 s |
| Lines found by speech-to-text | 8/8 (1 partial), 8/8, 14/16 (3 partial) | 8/8, 8/8, 16/16 |
| Output file | 16-bit stereo WAV, 44.1 kHz | Stereo MP3, 44.1 kHz, 128 or 256 kbps |
| Loudness | −13.4 to −15.2 LUFS | −11.4 to −13.2 LUFS |
| Cost and status | ||
| Price basis in Fuser | Per second of the Duration cap | Flat per song |
| At MiniMax | In the current API reference; open weights | Not in the current API reference |
They do different jobs. Music 3 follows a written style prompt and section tags and can run up to five minutes; Music-01 copies the sound of a reference song, and MiniMax stated up to 60 seconds for it. On the same lyrics, Music 3 produced longer files and Music-01's length tracked the number of lines. We measured length, lines sung, format and loudness, not musical quality.
Not in Fuser. The Music 3 node has no audio input; style comes from the text prompt. To reuse a sound, generate with Music 3, then pass that track to Music-01 as its reference.
MiniMax states up to five minutes for Music 3, and Fuser's Duration control goes to 300 seconds. MiniMax stated up to 60 seconds for Music-01, and our longest Music-01 track, from a 498-character lyric, was 57.1 seconds.
The song is cut off. With a 16-line lyric and a 30-second cap, the transcript found six lines before the file ended at 30.0 seconds. The same lyric with a 120-second cap ran 61.1 seconds.
Music-01 is a flat price per song. Music 3 is priced per second of the Duration you set, so a 60-second cap costs about 3.4 times a Music-01 song and a 30-second cap about 1.7 times.
In our runs Music 3 returned 16-bit stereo WAV at 44.1 kHz, and Music-01 returned stereo MP3 at 44.1 kHz, at 256 kbps in two runs and 128 kbps in one.
Describe a song for Music 3, reuse its sound with Music-01, and send either track straight into video.