# Lyria 3.5 Prompt Guide: Genre, Lyrics, Structure and Length Canonical page: https://fuser.studio/articles/lyria-3-5-prompt-guide How to prompt Google Lyria 3.5: genre, BPM, voice, your own lyrics, timestamps and length hints, tested on eleven real tracks with measured results. [All guides](https://fuser.studio/articles) · [Lyria 3.5 in Fuser](https://fuser.studio/models/lyria-3-5) **Quick answer:** Lyria 3.5 in Fuser takes one text prompt (up to 5,000 characters) and one optional image, and returns an MP3 song plus a Lyrics text output. Everything you control goes in the prompt: lead with the genre, give the tempo as a number such as "82 BPM", name the key, the instruments, the mood and the voice, then either paste your own words under a "Lyrics:" header with [Verse 1] and [Chorus] tags or write "Instrumental only, no vocals". Length is a prompt hint as well. In our runs "a 2-minute track" came back at 117.6 seconds and "a 60-second track" at 66.0 seconds, but three attempts at 30 seconds all came back at about 61 seconds, and prompts with no length ran from 1:59 to 2:57. ## What the Lyria 3.5 node takes and returns Lyria 3.5 is Google DeepMind's music model, announced on 29 July 2026 with better musicality, lyrics and vocals, and the promise that you can "more easily control the tempo and duration of your outputs" ([Google](https://blog.google/innovation-and-ai/models-and-research/google-labs/lyria-3-5/)). Google's model card lists its outputs as "Audio (music), text (lyrics)" ([model card](https://deepmind.google/models/model-cards/lyria-3-5/)). Google's API documentation describes it as the model for "full-length songs with verses, choruses, bridges", lasting "a couple of minutes", in 44.1 kHz stereo ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). It is a different model from Google's 30-second Lyria 3 Clip, which Fuser does not run. The [Lyria 3.5 node](https://fuser.studio/models/lyria-3-5) in Fuser has four sockets: - **Prompt** (input). Up to 5,000 characters. It is the only place to set style, structure, words and length. - **Image Inspiration** (input, optional). One image. Google's API accepts up to 10 images per request ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)); the Fuser node passes one. - **Generated Song** (output). An MP3. Every file in our tests was 44.1 kHz stereo at about 192 kbps. - **Lyrics** (output). A text version of the song: your lyrics or the ones the model wrote, with section labels (more on the format below). There is no duration, seed, tempo or negative-prompt setting. Google notes that "results may vary between calls, even with the same prompt" ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)), and with no seed you cannot pin a result, so plan on generating a few versions. Each run costs the same flat amount whether the song lasts one minute or three. All tracks in this guide were generated with the same model version the Fuser node runs: twelve runs of eleven different prompts (one prompt ran twice), which gave eleven tracks and one refusal. We measured them rather than describing how they sound: length with ffprobe, loudness per second, a simple beat-tracking estimate for tempo, and a speech-to-text pass to check which words were sung. ## Build the prompt in this order Google's own prompt guide says to "lead your prompt with the primary genre" and adds that "both short and detailed prompts produce strong results" ([Lyria prompt guide](https://ai.google.dev/gemini-api/docs/lyria-prompt-guide)). This is the order we used for every test: 1. **Genre.** One clear primary genre first: "Lo-fi hip hop", "Indie pop", "Cinematic orchestral score". 2. **Tempo and key.** Write them as numbers and names: "82 BPM in F major". Google's examples use exactly this form, including "slow tempo around 72 BPM" ([Lyria prompt guide](https://ai.google.dev/gemini-api/docs/lyria-prompt-guide)). 3. **Instruments.** Name them. Google: "If you want specific instruments or unusual combinations, declare them explicitly." 4. **Mood.** A few adjectives. Google's list includes Chill, Dreamy, Euphoric, Melancholic, Nostalgic and Triumphant. 5. **Voice, or no voice.** Describe the singer ("Female alto lead vocal, warm and close-miked") or write "Instrumental only, no vocals", the wording Google documents for instrumentals ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). Google's guide describes voice types such as "Female Alto: Rich, warm, and husky lower range" and "Male Baritone: Deep, velvet-smooth chest voice." 6. **Structure and length.** Section order, timestamps and a length hint go last. Our lo-fi test prompt. The six length runs below changed only its last sentence: ``` Lo-fi hip hop, instrumental only, no vocals. 82 BPM in F major. Dusty vinyl crackle, mellow Rhodes piano chords, a relaxed boom-bap drum groove and a warm upright bass. Calm, nostalgic, late-night mood. A 60-second track. ``` **Tempo held in most runs.** Across the six lo-fi runs at a requested 82 BPM, our beat estimate landed between 81.5 and 83.5 BPM in four. A fifth read 162.5, which is the same pulse counted at double time (about 81 BPM), and the sixth was ambiguous. The Spanish pop track asked for 100 BPM and measured 100. On the indie, soul and orchestral tracks our estimator could not settle on a tempo, so we make no claim about those. **"Instrumental only, no vocals" held.** We ran speech-to-text over three of the instrumental tracks (lo-fi, orchestral and the image test). It found no lyric lines in any of them; the only text was a single stray word on the image track, the kind of word speech-to-text tends to invent over music. ## Lyrics: paste your own or let the model write them Google's format for your own words is a "Lyrics:" header followed by section tags, with parentheses for "backing vocals, echoes, or ad-libs" ([Lyria prompt guide](https://ai.google.dev/gemini-api/docs/lyria-prompt-guide)), and its API docs ask you to keep lyrics clearly separate from the musical direction ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). This is the prompt we sent, shortened here: ``` Indie pop, 118 BPM in A major. Bright jangly electric guitar, punchy live drums, round bass, handclaps on the chorus. Female alto lead vocal, warm and close-miked. Structure: [Verse 1] -> [Chorus] -> [Verse 2] -> [Chorus] -> [Outro]. The chorus kicks in at 22 seconds. About 90 seconds long. Lyrics: [Verse 1] Paper lanterns on the fire escape Coffee going cold beside the window ... [Chorus] Stay up, stay up (stay up) Till the streetlights fade to gold ... [Outro] Fade to gold (fade to gold) ``` The Lyrics output came back with all 17 of our lines, word for word and in order, including the parenthesised echoes. Speech-to-text heard the verse lines and the "Stay up, stay up" chorus in the audio, so the words were sung, not just returned as text. Whether the parenthesised echoes were sung as a separate backing part is something to judge by ear; speech-to-text did not pick them out. ![Waveform of the first 45 seconds of a Lyria 3.5 indie-pop song with a moving playhead, above the verse and chorus lyrics we supplied.](https://statics.fuser.studio/cms/23f13d60-efea-40bf-a625-070b09703acd) _The first 45 seconds of the indie-pop test, sung from our lyrics. Open the clip to play it with sound._ **Letting the model write.** Leave out the Lyrics block and describe the subject instead. Our soul prompt asked for "a song about driving home at dawn after a long night shift" with "Verse, chorus, verse, chorus, bridge, final chorus". Lyria wrote 30 lines and the section labels in its Lyrics output followed that exact order. Speech-to-text confirmed the words were sung. **Language follows the prompt.** Google states that Lyria 3.5 "generates lyrics in the language of your prompt" ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). We wrote a pop prompt entirely in Spanish, with no lyrics, and got Spanish lyrics back; speech-to-text identified the vocal as Spanish. **What gets refused.** Google says prompts that "request specific artist voices or the generation of copyrighted lyrics" are blocked ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). Our one test that asked for a named pop star's voice returned a content-policy error and no audio. Paste only lyrics you wrote or have the rights to. For a deeper look at verse, chorus and bridge writing across music models, see [how to structure song lyrics for AI](https://fuser.studio/articles/ai-song-lyrics-structure-guide). ## Length: what a duration hint actually does Google says the model makes songs of "a couple of minutes" and that "exact duration can be influenced through your prompt", for example "create a 2-minute song" ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). We ran the same lo-fi prompt with different length requests: ![Waveforms of six Lyria 3.5 tracks from the same lo-fi prompt: no length 171.0 s, 2 minutes 117.6 s, 60 seconds 66.0 s, and three 30-second requests at 60.7, 61.1 and 61.6 s.](https://statics.fuser.studio/cms/400f161d-c3aa-45d9-b94e-8ff70a8b7450) _Six runs of the same lo-fi prompt with different length requests. The orange line marks the length asked for._ - **No length given:** 171.0 s. Our other prompts without a length hint came back at 167.1 s (soul), 176.7 s (image) and 118.6 s (Spanish pop). - **"A 2-minute track":** 117.6 s. - **"A 60-second track":** 66.0 s. - **"A 30-second track":** 60.7 s, then 61.1 s when we ran the identical prompt again. - **Timestamps ending at 0:30** plus "The track ends at 0:30": 61.6 s. So the hint moved the length at one and two minutes, but three different ways of asking for 30 seconds all produced a song of about a minute. Those are three runs, not a rule, and Google does not publish a minimum. If you need a short cue, generate a minute and cut it in your editor. ## Timestamps and section cues Google's guide says you can "prompt specific timing markers", such as "The chorus kicks in at 22 seconds", and its API docs show the [0:00 - 0:10] format ([Lyria prompt guide](https://ai.google.dev/gemini-api/docs/lyria-prompt-guide), [Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). We tested both. ![Two Lyria 3.5 waveforms with requested sections shaded: an orchestral cue of 105.2 s against sections ending at 1:05, and a song whose first detected chorus line is at 32.5 s.](https://statics.fuser.studio/cms/37e0f086-e481-41d1-8d20-b00f9c33969f) _Requested sections (blue) against the audio. Orange bars are sung lines found by speech-to-text, which can miss lines._ **Orchestral cue with four timestamped sections** (quiet cello intro to 0:15, a build to 0:35, a loud peak to 0:55, then a single piano note to 1:05). The first two changes landed close to the marks: the level stepped up at 16 seconds and again at 34 seconds. The rest did not: the "peak" was no louder than what followed it, there was only a one-second dip near 1:05 instead of a solo piano outro, and full-level music carried on to about 1:38 before fading out at 1:43. **Indie pop with "the chorus kicks in at 22 seconds" and "about 90 seconds long".** The first chorus line speech-to-text picked up was at 32.5 seconds, not 22 (it can miss lines, so the chorus may have started a little earlier). The song itself ended at about 94 seconds, close to the 90 we asked for, but the file ran to 118.6 seconds: 12 seconds of silence, then a quiet tail of about 12 seconds. Treat timestamps as a description of order and proportion, not a cue sheet. Check the waveform on the node before you cut to it, and expect to trim the end. ## Using an image as inspiration Connect any image to the Image Inspiration socket. Google says the model "will compose music inspired by the visual content" ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). We connected the still of a clay vase that we also use as a test input in our video guides, with only "Instrumental only, no vocals. Music inspired by this image." as the prompt. It returned a 2:57 instrumental with five sections, and speech-to-text found no lyric lines. A measurement cannot tell you whether the mood suits the picture; listen to it against the image. If you need a particular genre or tempo, put it in the prompt alongside the image, since the prompt is the documented control for both. ## Reading the Lyrics output The Lyrics socket is useful even for instrumentals. In all eleven of our tracks it returned a section map: each section starts with a label such as [[A0]] or [[B1]], where the letter marks the kind of section and repeats when that section comes back, and the number counts up in order. In our own-lyrics song both verses were A and both choruses B; the soul song came back A, B, A, B, C, B, which is verse, chorus, verse, chorus, bridge, chorus. Sung lines follow their label, each prefixed with [:]. Instrumental tracks return the labels only: the two-minute lo-fi track was A, B, C, B, C, D. Because Lyrics is an output socket like the song, the words and the section map stay on the canvas next to the track, so you can see how many sections you got before you listen. ## Watermarking and use Google states that "all generated audio includes a SynthID audio watermark for identification", which it says "is imperceptible to the human ear" ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). The model card names SynthID among its safety measures and points to Google's Generative AI Prohibited Use Policy for what the model must not be used for ([model card](https://deepmind.google/models/model-cards/lyria-3-5/)). We did not test the watermark ourselves. ## Lyria 3.5 or MiniMax Music 3? Fuser also runs [MiniMax Music 3](https://fuser.studio/models/minimax-music-3), which works differently. Its node has a separate Lyrics field, a Duration setting from 1 to 300 seconds, guidance and inference-step controls, and a seed. Lyria 3.5 has none of those and puts everything in one prompt, but it accepts an image and returns the lyrics with a section map. On cost, Lyria charges one flat amount per song, while MiniMax Music 3 is priced by the duration you set: a MiniMax song set to 60 seconds costs about a fifth more than a Lyria song, the two cost about the same at 50 seconds, and shorter MiniMax songs cost less. We ran both on the same prompts in [Lyria 3.5 vs MiniMax Music 3](https://fuser.studio/articles/lyria-3-5-vs-minimax-music-3), and the [MiniMax Music 3 prompt guide](https://fuser.studio/articles/minimax-music-3-prompt-guide) covers its controls. For scoring picture, see [AI music for video](https://fuser.studio/articles/ai-music-for-video); for the wider field, see [the best AI music and sound effect generators](https://fuser.studio/articles/best-ai-music-and-sound-effect-generators). ## Lyria 3.5 cheat sheet. Each prompt technique, how to write it, and what it did in our runs. ### Prompt | Technique | How to write it | What we measured | | --- | --- | --- | | Genre | Lead with one primary genre. | 10 of our 11 prompts led with one; the image test did not. | | Tempo | A number: "82 BPM". | 81.5 to 83.5 BPM in 4 of 6 lo-fi runs, a 5th at double time; 100 asked, 100 measured on the Spanish track. | | Instrumental | "Instrumental only, no vocals". | No lyric lines found on the 3 tracks we transcribed. | | Own lyrics | "Lyrics:" header, [Verse 1], [Chorus], (backing). | All 17 lines returned in order; verse and chorus heard in the audio. | | Model-written lyrics | Describe the subject and the section order. | 30 lines, sections in the order asked. | | Language | Write the prompt in the language you want sung. | Spanish prompt, Spanish vocal. | ### Length and structure | Technique | How to write it | What we measured | | --- | --- | --- | | No length | Leave it out. | 118.6 to 176.7 s across four prompts. | | Length hint | "A 2-minute track". | 2 min: 117.6 s. 60 s: 66.0 s. 30 s: about 61 s, three times. | | Timestamps | [0:00 - 0:15] Intro: ... | First two section changes within a second of the marks; later ones ignored. | | Image | Connect to Image Inspiration. | 2:57 instrumental, five sections. | ## Questions, answered. ### Can Lyria 3.5 sing my own lyrics? Yes. Put them under a "Lyrics:" header with section tags such as [Verse 1] and [Chorus], separate from the style description. In our test the Lyrics output returned all 17 of our lines in order and speech-to-text heard the verse and chorus sung in the audio. ### How long are Lyria 3.5 songs? Google describes them as "a couple of minutes". Our prompts with no length hint ran from 1:59 to 2:57. "A 2-minute track" gave 117.6 seconds and "a 60-second track" 66.0 seconds, but three requests for 30 seconds all came back at about 61 seconds. ### Can I set the duration or seed for Lyria 3.5 in Fuser? No. The node has a prompt and one optional image input, and no duration, seed, tempo or negative-prompt setting. Length, tempo and structure are written into the prompt, and results vary between runs of the same prompt. ### How do I make an instrumental with Lyria 3.5? Write "Instrumental only, no vocals" in the prompt, the wording Google documents. Speech-to-text found no lyric lines on the three instrumental tracks we checked. ### Does Lyria 3.5 music have a watermark? Google states that all generated audio includes an imperceptible SynthID audio watermark for identification. ### What does the Lyrics output contain for an instrumental? Section labels only, such as [[A0]], [[B1]], [[C2]]. The letter repeats when a section returns, so the labels work as a map of the arrangement. ## Write the song in one prompt, then wire it to the picture. Run Lyria 3.5 next to your image and video models on one canvas, and keep the lyrics beside the track. [Open Lyria 3.5 in Fuser](https://fuser.studio/models/lyria-3-5) · [Explore all guides](https://fuser.studio/articles) ## More articles - [MiniMax Music Prompt Guide: Reference Songs, Lyric Formatting and Length](https://fuser.studio/articles/minimax-music-prompt-guide.md) - [MiniMax Music 3 Prompt Guide: Structured Captions, Section Tags and Song Length](https://fuser.studio/articles/minimax-music-3-prompt-guide.md) - [How to Structure Song Lyrics for AI: Verse, Chorus and Bridge Tags That Work](https://fuser.studio/articles/ai-song-lyrics-structure-guide.md)