# MiniMax Music 3 Prompt Guide: Structured Captions, Section Tags and Song Length Canonical page: https://fuser.studio/articles/minimax-music-3-prompt-guide How to prompt MiniMax Music 3: write a three-part Structured Caption, place section tags on their own lines, get a clean instrumental and set the duration ceiling, tested on 11 real songs. [All guides](https://fuser.studio/articles) · [MiniMax Music 3 in Fuser](https://fuser.studio/models/minimax-music-3) **Quick answer:** In Fuser, MiniMax Music 3 takes two required texts: a music description and lyrics. Write the description as MiniMax's three-part Structured Caption (Global Metadata, Vocal Details, Arrangement), put section tags such as [verse] and [chorus] on their own lines, and treat Duration as a ceiling, not a target. In two same-seed pairs a detailed caption got more of the lyric sung intelligibly than a one-line prompt, a lyric of only [instrumental] still produced wordless vocals while a tags-only lyric gave a clean instrumental, and a ceiling set too long filled the gap with made-up vocals. ## What MiniMax Music 3 is MiniMax released Music 3.0 on 13 August 2026 as an open-weights model that "composes, arranges, performs, and produces a complete song in a single generation", with songs up to five minutes long ([MiniMax announcement](https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model)). The official model card describes it as conditioned on "lyrics and a detailed music description" ([MiniMax-Music3 model card](https://huggingface.co/MiniMaxAI/MiniMax-Music3)). The weights are published under the MiniMax-Music3 Community License. Among its terms, a commercial product built on the weights must display "MiniMax-Music3" in its interface, needs separate written authorisation from MiniMax once yearly revenue exceeds 20 million US dollars, and its list of prohibited uses includes putting machine-generated content into any public environment without clearly disclosing that it is machine-generated ([licence](https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE)). Read it in full if you plan to publish the songs commercially. This guide covers Music 3 only. Fuser also runs MiniMax's older reference-song model, which works differently (a reference track plus lyrics, no style text); that one has its own [MiniMax Music prompt guide](https://fuser.studio/articles/minimax-music-prompt-guide), and [MiniMax Music 3 vs Music-01](https://fuser.studio/articles/minimax-music-3-vs-music-01) compares the two. ## The inputs, as Fuser exposes them - **Prompt (required).** The music description: style, mood, vocals, instruments, tempo and arrangement. - **Lyrics (required).** The words to sing, one line per line, with section tags. MiniMax's announcement describes lyrics as optional for the model itself, but Fuser's node has no instrumental switch and will not run with an empty lyrics box (see the instrumental section below for what to type instead). - **Duration.** 1 to 300 seconds, default 60. It is an upper bound: the model can stop earlier, and the node returns the real length as a separate Actual Duration output. - **Inference Steps.** 1 to 100, default 30. These are flow-matching steps per eight-second chunk of audio, so more steps cost generation time on every chunk. - **Guidance.** 0 to 20, default 1.7: how strongly generation follows the description. - **Seed.** Reuse it to reproduce a result. The output is a WAV file. MiniMax's model card says the model "produces 32 kHz, 16-bit stereo WAV audio"; every file we received (11 of them) measured 44.1 kHz, 16-bit, stereo. We report what we measured. ## Step 1: write the prompt as a Structured Caption MiniMax's model card recommends "a Structured Caption with three sections" for precise control ([model card](https://huggingface.co/MiniMaxAI/MiniMax-Music3)): - **Global Metadata:** genre, subgenre, BPM, key, scale, emotional progression, listening scenario and production profile. - **Vocal Details:** vocal gender, timbre, performance style, harmony, backing vocals and vocal effects. - **Arrangement:** primary and secondary instruments, how instruments change section by section, groove, bass, percussion, textures and spatial effects. The same card says a concise natural-language description can be used directly; the Structured Caption is the route it gives for richer, more precise control. MiniMax's announcement lists the same ground in plainer words, including tempo, time signature, key, the "emotional contour", when instruments enter and leave, and "section-level changes in vocal delivery" ([announcement](https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model)). Here is the caption we used, written to that template (line breaks added for reading; we sent each of the three sections as a single line): ``` Global Metadata: Genre: indie pop. Subgenre: synth-driven bedroom pop. BPM: 104. Key: A major. Emotional progression: wistful in the verses, lifting into a bright, open chorus, hushed bridge, full final chorus. Listening scenario: late-night city drive. Production profile: warm analog synths, tight modern mix, wide stereo. Vocal Details: Female lead, clear and airy, slightly breathy. Intimate in the verses, belted in the chorus. Two-part harmonies in the chorus. Soft "ooh" backing vocals in the pre-chorus. Light plate reverb, slap delay on line ends. Arrangement: Primary instruments: analog synth pads and electric piano. Secondary: clean electric guitar, synth bass. Intro is pads and a filtered drum loop; drums and bass enter in verse 1; guitar arpeggio joins in the pre-chorus; full kit with open hi-hats in the chorus; bridge drops to electric piano and voice only; final chorus adds a synth lead. Groove: steady four-on-the-floor kick with syncopated hi-hats. Bass: round synth bass on the root. Percussion: claps on 2 and 4 in the chorus. Spatial: wide pads, centred vocal. ``` To see whether the detail matters, we ran the same lyric twice per seed: once with this caption, once with "An upbeat indie pop song with female vocals." Seed 4242 used the full 16-line lyric (26 sung lines with repeats) and a 150-second ceiling; seed 777 used only the intro, first verse, pre-chorus and chorus with a 90-second ceiling. ![Waveforms of two 150-second MiniMax Music 3 songs with the same lyrics and seed; the Structured Caption version shows many more transcribed lyric lines than the one-line prompt version.](https://statics.fuser.studio/cms/c2bedc43-863d-4e00-bbbd-3a8e1847e919) _Same lyrics, seed 4242, 150-second ceiling, default steps and guidance. Coloured bars are lines a speech-to-text pass could match to our lyric; it can miss sung lines, so treat them as a lower bound. Generated 2 October 2026._ - **Lyric delivery.** A speech-to-text pass matched 13 of the 16 distinct lines in the seed-4242 caption version against 5 in the one-line version, and 6 of 10 against 4 of 10 on seed 777. In both one-line versions it found none of the first verse. A second check with an audio-understanding model agreed on both pairs: the one-line versions replaced the first verse with garbled, unintelligible words, while the caption versions sang it. - **Tempo.** The caption asked for 104 BPM. Beat tracking put both seed-4242 songs at about 103, and the seed-777 pair at about 115 (caption) and 99 (one-line). Two pairs can't show that the BPM field is followed, so check tempo by ear if it matters. - **What we can't claim.** Two pairs on two seeds is a consistent pattern, not a benchmark. MiniMax's own wording is that tags and descriptions "provide generative control rather than strict symbolic guarantees". The practical reading: spell out the vocal and the arrangement, section by section. The model has more to work with, and in our runs it held onto the words better. ## Step 2: tag the sections, each tag on its own line The model card lists nine structure tags: [Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Instrumental], [Solo] and [Outro], and says to put them "on their own lines" ([model card](https://huggingface.co/MiniMaxAI/MiniMax-Music3)). MiniMax's announcement names the same set minus [Post-Chorus] ([announcement](https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model)). Our lyrics used lowercase forms ([verse], [pre-chorus]), and the transcribed lines came back in lyric order. A tagged lyric looks like this: ``` [intro] [verse] Streetlights blinking out of time Radio is losing the line [pre-chorus] And I keep the window down [chorus] We were neon on the overpass ``` Text on the same line as a tag is lost. We ran a short rock lyric with the first verse line and the first chorus line written straight after their tags ("[verse] Paper planes above the playground"), then the identical lyric with the tags on their own lines, same seed and settings. In the same-line version both of those lines were missing, confirmed by speech-to-text and by the audio-understanding check; the model sang "We were faster than the bell" once, "Run until the streetlights tell" four times, and filled the gaps with unintelligible words. In the own-line version all four lines were sung, in order. You don't have to police this in Fuser. The MiniMax Music 3 node moves any text that follows one of those nine tags onto the next line before sending the lyric, so "[chorus] Run until the summer's over" arrives as a tag line plus a lyric line. It only does this for those nine tags; any other bracketed word is passed through as typed. ## Step 3: getting an instrumental There is no instrumental switch, and the lyrics box can't be empty. We tried two ways round it with the same ambient caption (felt piano, string pad, synth bass, brushed percussion, "No vocals. Purely instrumental."), seed 4242 and a 60-second ceiling: - **Lyrics of only [instrumental]:** the track ran the full 60.07 seconds and was not instrumental. Speech-to-text picked up repeated syllables from about 31 seconds on, and the audio-understanding check heard female humming and "che, che, che" syllables over solo piano. - **Lyrics of [intro], [instrumental], [solo] and [outro], each on its own line:** the track stopped on its own at 48.09 seconds, short of the 60-second ceiling. Speech-to-text found no words, and two separate passes with the audio-understanding model heard no voice. Both described piano and a pad, with bass, drums and a flute-like lead coming in around the middle; we hadn't asked for the flute. One run each, so this is not a guarantee, but a lyric made only of non-vocal section tags is the version that worked. ## Step 4: set Duration as a ceiling that fits the lyric ![Five MiniMax Music 3 tracks on a timeline showing requested duration ceilings against actual lengths, with sung lyric lines and unintelligible vocals marked.](https://statics.fuser.studio/cms/4ad1099f-4d67-413d-94d2-2b03450ac247) _Duration is a ceiling, not a target. Orange: lines from our lyric. Grey hatching: vocals that matched no lyric. Generated 2 October 2026 with the same model version Fuser runs._ Every run either filled the ceiling or stopped short of it, and both ends have a failure mode: - **Ceiling too long for the lyric.** Four lines with a 150-second ceiling stopped at 115.07 seconds. The four lines were sung once, then speech-to-text found stretches of vocals between about 55 and 91 seconds that matched no line of our lyric, before the track faded out. The same four lines with a 40-second ceiling ran 40.04 seconds, with all four lines sung and about four seconds of unintelligible vocal at the end. - **Ceiling too short for the lyric.** The full 26-line song with a 150-second ceiling ran exactly 150.19 seconds and was still at full level in the final second, about 34 dB louder than the natural fade at the end of the 115-second track: it was cut off, not finished. With a 240-second ceiling, the same prompt, lyric and seed ended on its own at 156.44 seconds, with a near-silent final second (about -52 dB). That lyric needed about 156 seconds; 150 was six seconds short. - **Spare time gets filled with words you didn't write.** On seed 777 the shorter lyric (one verse, pre-chorus and chorus) filled the whole 90-second ceiling, and the audio-understanding check heard extra lines that weren't in our lyric after the chorus in both versions. As a working rule from these runs: our complete song took about six seconds per sung line once the intro, breaks and outro were included (156 seconds for 26 lines). Use that to size the ceiling, then check the Actual Duration output. If it equals the ceiling and the end is loud, raise the ceiling. If the song runs on with vocals that aren't your words, lower it or give the model more lyric. Cost follows the ceiling, not the result: Fuser estimates and bills Music 3 by the requested duration, so a 300-second ceiling on a 60-second idea pays for five minutes. A 50-second ceiling costs about the same as one [Lyria 3.5](https://fuser.studio/models/lyria-3-5) song, a 30-second ceiling a little over half of one, and the full five minutes about six times as much. ## Seeds, steps and guidance - **Seed.** We re-ran the 40-second rock song with identical inputs and seed: the two files were sample-for-sample identical. Keep the seed when you want to change only one thing. - **The ceiling is part of the recipe.** Same lyric, prompt and seed with a different ceiling gave different audio in both of our comparisons, though the layout held: the rock vocal entered at about 14.6 seconds with a 40- or 150-second ceiling, and the full song's second verse started at about 61.5 seconds with a 150- or 240-second ceiling. Changing Duration renders a new take, not a trim. - **Steps and guidance.** We left both at the defaults (30 and 1.7) for every run, so we have no measurements to offer on them. Per the documentation, steps apply to each eight-second chunk and guidance controls how closely the generation follows the description. ## Common problems and fixes - **Words come out as nonsense.** Add detail to the Vocal Details and Arrangement parts of the caption; a one-line prompt lost most of our first verse. - **A line is missing.** Check it isn't on the same line as a tag (Fuser fixes the nine standard tags for you) and that the song wasn't cut off by the ceiling. - **Track ends abruptly.** Actual Duration equals the ceiling: raise it. - **Extra vocals after the lyric.** The ceiling is longer than the lyric needs: lower it, or add an [outro] and more lines. - **Vocals in an instrumental.** Don't rely on [instrumental] alone; use non-vocal tags only, such as [intro], [instrumental], [solo], [outro]. - **Need the same song again.** Reuse the seed and every other setting, including Duration. ## Build it as a workflow In Fuser the prompt and lyrics can live in their own Text nodes wired into MiniMax Music 3, so you can swap a caption without retyping the lyric. The Generated Song output connects to other nodes: [Wan 2.6 Video](https://fuser.studio/models/wan-2-6-video) takes audio as background music and keeps the first 5, 10 or 15 seconds to match the clip; [LatentSync](https://fuser.studio/models/latentsync) lip-syncs a video to it; [Whisper](https://fuser.studio/models/whisper) transcribes it for captions. For scoring a cut end to end, see [AI music for video](https://fuser.studio/articles/ai-music-for-video); for writing the lyric itself, the [song lyrics structure guide](https://fuser.studio/articles/ai-song-lyrics-structure-guide). If you are choosing between models, [Lyria 3.5 vs MiniMax Music 3](https://fuser.studio/articles/lyria-3-5-vs-minimax-music-3) and the [Lyria 3.5 prompt guide](https://fuser.studio/articles/lyria-3-5-prompt-guide) cover Google's model, and [the best AI music and sound effect generators](https://fuser.studio/articles/best-ai-music-and-sound-effect-generators) covers the rest of Fuser's audio lineup. ![Waveform with a moving playhead for a 45-second excerpt of the Structured Caption song generated with MiniMax Music 3.](https://statics.fuser.studio/cms/ed43b740-59df-49e1-a108-f115692bb4e1) _Seconds 18 to 63 of the Structured Caption song from the 240-second-ceiling run, ending as the second verse begins. Open the clip to play it with sound._ ## MiniMax Music 3 cheat sheet. Each control, how to use it, and what it did in our runs. ### Inputs | Control | How to use it | What we measured | | --- | --- | --- | | Prompt | Structured Caption: Global Metadata, Vocal Details, Arrangement. | More lines sung clearly than a one-line prompt: 13 vs 5 of 16, and 6 vs 4 of 10. | | Lyrics | Required. One sung line per line; tags on their own lines. | Text after a tag on the same line was not sung. | | Instrumental | Lyric of non-vocal tags only: [intro] [instrumental] [solo] [outro]. | 48.1 s, no voice found. [instrumental] alone: vocals. | | Duration | Ceiling, 1 to 300 s. Size it from the lyric, then check Actual Duration. | 26 lines: cut off at 150 s; ended on its own at 156 s of 240. | | Seed | Reuse with every other setting unchanged. | Identical inputs and seed: identical audio. | | Steps / Guidance | Defaults 30 and 1.7. | Left at defaults in all runs. | ## Questions, answered. ### How do I prompt MiniMax Music 3? Write the music description as a Structured Caption in three parts: Global Metadata (genre, BPM, key, emotional progression), Vocal Details (gender, timbre, delivery, harmonies, effects) and Arrangement (instruments and how they change by section). Put the lyrics in the separate Lyrics field with section tags on their own lines. ### What section tags does MiniMax Music 3 support? MiniMax's model card lists [Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Instrumental], [Solo] and [Outro]. Each tag goes on its own line; text written on the same line as a tag was not sung in our test. ### Can MiniMax Music 3 make instrumental music? In Fuser there is no instrumental switch and the Lyrics field is required. A lyric of only [instrumental] still produced wordless vocals in our run; a lyric of [intro], [instrumental], [solo] and [outro] on separate lines gave a 48-second track with no voice detected. ### How long can a MiniMax Music 3 song be? Up to five minutes. In Fuser, Duration is a ceiling from 1 to 300 seconds; the model can stop earlier and reports the actual length. Set it a little above what the lyric needs: too long and the model adds vocals that are not your words, too short and the song is cut off. ### What audio format does MiniMax Music 3 return? A WAV file. MiniMax's model card states 32 kHz, 16-bit stereo; the 11 files we generated measured 44.1 kHz, 16-bit stereo. ### Is MiniMax Music 3 open source? The weights are open and published under the MiniMax-Music3 Community License, which sets conditions on commercial products built on them and prohibits putting machine-generated content into public environments without clearly disclosing it. ## Caption, lyric, ceiling. Then wire the song to the picture. Generate full songs with MiniMax Music 3 and send them straight into video, lip sync and captions on one canvas. [Open MiniMax Music 3 in Fuser](https://fuser.studio/models/minimax-music-3) · [Explore all guides](https://fuser.studio/articles) ## More articles - [MiniMax Music Prompt Guide: Reference Songs, Lyric Formatting and Length](https://fuser.studio/articles/minimax-music-prompt-guide.md) - [MiniMax Music 3 vs Music-01: Same Lyrics, Measured Side by Side](https://fuser.studio/articles/minimax-music-3-vs-music-01.md) - [How to Structure Song Lyrics for AI: Verse, Chorus and Bridge Tags That Work](https://fuser.studio/articles/ai-song-lyrics-structure-guide.md)