# AI Music for Video: Score a Generated Clip with Lyria 3.5 or MiniMax Music 3 Canonical page: https://fuser.studio/articles/ai-music-for-video How to score a generated video clip with MiniMax Music 3 or Lyria 3.5: matching length and mood, joining picture and music, and using your track as reference audio. Tested. [All guides](https://fuser.studio/articles) · [Lyria 3.5](https://fuser.studio/models/lyria-3-5) · [MiniMax Music 3](https://fuser.studio/models/minimax-music-3) **Quick answer:** To score a generated clip in Fuser, read the clip's length, then generate the music to fit. [MiniMax Music 3](https://fuser.studio/models/minimax-music-3) has a Duration setting: we set it to 10 for a 10-second clip and got a 10.00-second track. [Lyria 3.5](https://fuser.studio/models/lyria-3-5) takes the length only as a prompt hint plus one image for mood; "a 10-second track" came back at 62.5 seconds, so you cut a section from it. Fuser has no node that lays a music file under an existing clip, and the Compositor exports video without sound, so you join picture and music in your video editor after downloading both. If you want the music inside a generated video instead, [Wan 3.0](https://fuser.studio/models/wan-3-0-video) can take your track as Reference Audio: in our test its output soundtrack was our track, essentially unchanged. ## What Fuser can and can't do with picture and music Be clear about this before you start, because it decides the workflow: - **Generating the music:** yes. Lyria 3.5 and MiniMax Music 3 each output an audio file on the canvas, next to the video node it is meant for. - **Laying a track under a clip you already have:** no. No node in Fuser mixes a music file into an existing video. The Compositor works on images and video frames and exports an MP4 with no audio track, and the lip-sync nodes that take a video plus an audio file re-animate a face to match speech. That step happens in your video editor. - **Generating sound for a clip you already have:** yes, but built for effects. [MMAudio](https://fuser.studio/models/mmaudio) and [Mirelo SFX](https://fuser.studio/models/mirelo-sfx) take the video itself and return it with a generated track synchronized to the action. Mirelo is described as a sound-effects model; we did not test either for a music score. - **Video with sound made in one pass:** yes, from the video models that generate their own audio (Kling O3, Wan 3.0, Seedance 2.5, Veo 3.1). That sound is theirs, not your track. - **A new video built around your track:** yes, by connecting your music as a reference to Wan 3.0 or Seedance 2.5. You get a new shot, not your original clip with music added. The steps below cover the first route, then the two alternatives, all tested on the same 10-second shot of a terracotta vase. Each test was one run. ## Step 1: make the clip and note its exact length We animated our vase still with [Kling O3](https://fuser.studio/models/kling-o3) Standard, image-to-video, Duration 10, and a slow push-in prompt. The file came back at 10.04 seconds, 24 frames per second. That number is your target: the score should be at least as long as the clip, and ideally end where the picture ends. ## Step 2a: score it to length with MiniMax Music 3 MiniMax Music 3 is MiniMax's music model, released on 13 August 2026 with open weights and songs "up to five minutes" long ([MiniMax](https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model)). The Fuser node has a Prompt, a Lyrics field, and a **Duration** slider from 1 to 300 seconds. The node treats Duration as a maximum, "the model can finish earlier", and reports the real length on its **Actual Duration** output, so check that number rather than assuming. Two things to know for scoring: 1. **Lyrics are required in Fuser.** For an instrumental, put a structure tag on its own: we used [instrumental], one of the tags MiniMax lists alongside [intro], [verse] and [outro] ([Hugging Face model card](https://huggingface.co/MiniMaxAI/MiniMax-Music3)). Speech-to-text found no words in the result. 2. **Write the cue like a brief.** Our prompt gave genre, tempo, key, the mood over time, and the arrangement, and named the job: "background score for a 10-second product film of a terracotta vase in late-afternoon sunlight". ``` Genre: cinematic ambient. BPM: 72. Key: D major. Emotional progression: calm and warm, a slow gentle swell, then settling. Listening scenario: background score for a 10-second product film of a terracotta vase in late-afternoon sunlight. Production profile: clean, wide, premium. Vocal Details: No vocals. Purely instrumental. Arrangement: Felt piano motif over a soft string pad; a low cello note enters halfway; ends on a sustained piano chord. Lyrics: [instrumental] Duration: 10 ``` **What came back:** a 10.00-second WAV. A low note entered at 4.85 seconds, close to the "halfway" we asked for. The track did not wind down at the end, though: its last tenth of a second was still at full level, so it stops rather than resolves. Fade the last half second in your editor, or set Duration a few seconds longer than the clip and choose where to cut. Because MiniMax Music 3 is priced by the duration you set, a 10-second cue costs about a fifth of one Lyria 3.5 song. The [MiniMax Music 3 prompt guide](https://fuser.studio/articles/minimax-music-3-prompt-guide) covers the prompt format in depth. ![Waveforms on one time scale: the Kling O3 clip at 10.04 s, MiniMax Music 3 with Duration 10 at 10.00 s, and Lyria 3.5 asked for a 10-second track at 62.5 s.](https://statics.fuser.studio/cms/65b923e0-4146-4d1e-96b3-0d140bc7931e) _The clip and both scores on the same time scale. Only MiniMax Music 3 came back at the length we needed._ ## Step 2b: or let Lyria 3.5 take its mood from the picture Lyria 3.5 is Google DeepMind's music model, announced on 29 July 2026 ([Google](https://blog.google/innovation-and-ai/models-and-research/google-labs/lyria-3-5/)). Google describes its output as songs of "a couple of minutes (controllable using prompt)", with duration set "by specifying it in your prompt" or by timestamps, in 44.1 kHz stereo ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). There is no duration setting. What it adds for video work is the **Image Inspiration** input: Google says the model composes music inspired by the images you give it. The API takes up to 10 images; the Fuser node takes one. We connected the vase still the clip started from to Image Inspiration and used the same brief in Lyria's style, ending with "A 10-second track." The song came back at **62.5 seconds**. It opened very quietly and built up: the first second measured about -50 dB and the music reached full level only around 6 seconds in, then ran steadily until a fade near the minute mark. For a 10-second clip that means picking a 10-second section, not using the start. Our [Lyria 3.5 prompt guide](https://fuser.studio/articles/lyria-3-5-prompt-guide) found the same pattern: a 60-second request returned 66 seconds and a 2-minute request 118 seconds, but all three 30-second requests came back at about a minute. Lyria suits a scene that needs a longer cue, or when you want the picture itself to steer the mood; MiniMax Music 3 suits a cue that must be an exact length. [Lyria 3.5 vs MiniMax Music 3](https://fuser.studio/articles/lyria-3-5-vs-minimax-music-3) compares them on the same prompts. Google states that all Lyria output carries an inaudible SynthID watermark ([Gemini API docs](https://ai.google.dev/gemini-api/docs/music-generation)). ## Step 3: put picture and music together in your editor Use **Download media** on the video node and on the music node, then in any video editor: 1. Place the music at the start of the timeline, under the clip. 2. Trim it to the clip's length (or slide a longer Lyria track until the section you want lines up). 3. Fade the last half second. 4. Mute the clip's own audio, or keep it under the music if it carries useful room sound. We did exactly this with the Kling clip and the MiniMax track, adding a 0.6-second fade. The track was 0.04 seconds shorter than the clip, too little to see. ![The Kling O3 vase clip with the MiniMax Music 3 track laid underneath in a video editor.](https://statics.fuser.studio/cms/d359cc96-76bf-44a5-9c19-fcfe32b1e7b4) _The Kling O3 clip with the MiniMax Music 3 score added in an editor. Inline playback is muted; open the clip to hear it._ ## Alternative 1: let the video model make its own sound Several video models in Fuser generate audio with the picture. It is quicker than scoring, but you don't choose the track, and you can't change the music later without regenerating the shot. - **Kling O3.** Kling describes its native audio as "synchronized dialogue, ambient sounds, and sound effects" ([Kling](https://kling.ai/blog/kling-video-3-omni-multi-shot-native-audio-guide)). In Fuser, **Generate Audio** is off by default. Our clip in Step 1 had it on, and the prompt ended with "Audio: a slow, warm solo piano melody over soft room tone". The soundtrack did contain music: a run of struck, decaying pitched notes from 0.3 seconds to the end, with no speech. It was also quiet, about 9 dB lower on average than the MiniMax track. - **Wan 3.0.** Alibaba says it "natively outputs dialogue, BGM, and sound effects" ([Alibaba Cloud](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-guide)). Its Generate Audio toggle is on by default in Fuser. - **Seedance 2.5** generates synchronized sound effects, ambience and speech when Generate Audio is on. - **Veo 3.1** is, in Google's words, "a model for generating video with native audio" ([Gemini API docs](https://ai.google.dev/gemini-api/docs/video)). ![The same Kling O3 vase clip with its own generated audio, prompted for a slow solo piano melody over room tone.](https://statics.fuser.studio/cms/79958aca-9242-463a-94ac-cbd66534d1b3) _The same clip with Kling O3's own audio, prompted for a slow piano melody. Open the clip to hear it._ For a side-by-side of these models' audio, see [the best AI video generators with audio](https://fuser.studio/articles/best-ai-video-generators-with-audio). ## Alternative 2: give the video model your track as a reference Two video models in Fuser accept audio as an input. Here the music comes first and the model generates a new video around it, so this suits a shot you haven't made yet. We sent each one the same three things: the vase still as a reference image, the 10-second MiniMax track as reference audio, and a prompt ending "Use Audio 1 as the background music for the whole video" (written "@Audio1" for Seedance). Both were set to 10 seconds with audio on. To check what came back, we compared each output soundtrack with our track sample by sample. As a control, the same comparison against the Kling clip's unrelated audio scored 0.02 overall. ![Three waveforms: the MiniMax Music 3 track, the Wan 3.0 output soundtrack matching it exactly, and the Seedance 2.5 soundtrack with the low note arriving at 5.5 s instead of 4.85 s.](https://statics.fuser.studio/cms/8b26d8ed-5b10-45cf-b660-c7f421dff928) _Our track and the two output soundtracks. Wan 3.0 kept it note for note; Seedance 2.5 kept the notes but stretched them out by about 12 to 14 percent._ **Wan 3.0: our track, essentially unchanged.** The [Wan 3.0 node](https://fuser.studio/models/wan-3-0-video) takes up to five Reference Audio clips totalling 15 seconds, addressed in the prompt as Audio 1, Audio 2 and so on ([Alibaba Cloud](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-api-reference)). Alibaba's guide describes audio references as a way to "provide music or voice to generate dance or lip-sync animation synchronized with the beat/tone" ([Alibaba Cloud](https://www.alibabacloud.com/help/en/model-studio/wan3-video-generation-guide)). In our 720p run on the Standard model, the output's soundtrack matched our track with a correlation of 0.98 at zero offset, within 1 dB of its level, with the low note at 4.85 seconds as in the original. The picture was a new 10.03-second shot of the vase, not our Kling clip. One limit: reference media can't be combined with a start or end image. Fuser asks you to remove them, so you can't pin the exact first frame in this mode. The [Wan 3.0 guide](https://fuser.studio/articles/wan-3-0-guide) covers the other reference types. ![A new Wan 3.0 shot of the terracotta vase on its plinth, generated with our MiniMax Music 3 track as Reference Audio.](https://statics.fuser.studio/cms/0cd1490e-ce92-4309-89cf-f77df55de86f) _Wan 3.0 with our track as Reference Audio: a new 10-second shot whose soundtrack is that track. Open the clip to hear it._ **Seedance 2.5: the right notes, stretched out.** The [Seedance node](https://fuser.studio/models/seedance-2) takes reference audio on Seedance 2.5 only: up to 10 MP3 or WAV files, each 2 to 30 seconds, 30 seconds in total. Fuser also needs at least one reference image or video alongside it, referred to as @Image1, @Audio1 and so on. Our 480p run kept our notes at the same pitch but stretched them out, so the music took about 12 to 14 percent longer to play. The low note arrived at 5.5 seconds instead of 4.85, and the gap grew as the track went on: under 0.1 seconds at the start, about 0.6 seconds by the 4-second mark and 1.2 seconds by 8 seconds. The 10.1-second video ran out when it had reached roughly 8.7 seconds into our track, so the ending we wrote never played. For a wider comparison of the two models, see [Wan 3.0 vs Seedance 2.5](https://fuser.studio/articles/wan-3-0-vs-seedance-2-5). These are single runs, so treat them as what happened once, not a guarantee. If the exact track matters, as it does for a licensed cue or a voiceover, check the waveform after generation, or keep the music out of the video model and combine the two in your editor as in Step 3. ## Choosing a route - **You have the clip and need music that fits it exactly:** MiniMax Music 3 with Duration set to the clip length, then your editor. - **You want the picture to set the mood, or need a longer cue:** Lyria 3.5 with the clip's start image as Image Inspiration, then cut a section in your editor. - **You haven't made the shot yet and the music is fixed:** connect the track to Wan 3.0 as Reference Audio. - **You want sound effects and ambience more than a score:** turn on the video model's own audio. For sound effects on their own, see the [ElevenLabs sound effects guide](https://fuser.studio/articles/elevenlabs-sound-effects-guide). All of these run as nodes on one canvas, so the clip, both scores and any reference-audio test stay side by side while you choose. For longer chains that take a still through video and sound, see [how to chain AI models in one workflow](https://fuser.studio/articles/chain-ai-models-workflow). ## Four ways to get music onto a clip. What each route does, and what it did in our tests on one 10-second shot. ### Score your clip, combine in an editor | Route | How it works in Fuser | What we measured | | --- | --- | --- | | MiniMax Music 3 | Duration slider (1 to 300 s, a maximum); Lyrics required, [instrumental] for no vocals. | Duration 10 gave 10.00 s; no words detected; ends at full level, so add a fade. | | Lyria 3.5 | Length only as a prompt hint; one Image Inspiration input. | "A 10-second track" gave 62.5 s; full level only from about 6 s. | ### Music made inside the video | Route | How it works in Fuser | What we measured | | --- | --- | --- | | Native audio | Generate Audio on Kling O3 (off by default), Wan 3.0 and Seedance 2.5 (on by default), Veo 3.1. | Kling O3 asked for piano: pitched notes throughout, about 9 dB quieter than our score. | | Wan 3.0 reference audio | Up to 5 clips, 15 s total; no start or end image in this mode. | Soundtrack matched our track: correlation 0.98, no offset. | | Seedance 2.5 reference audio | Up to 10 files, 30 s total; needs an image or video reference too. | Same notes and pitch, stretched about 12 to 14 percent; the last second or so of our track never played. | ## Questions, answered. ### Can Fuser add music to a video I already have? Not directly. No node mixes a music file into an existing video, and the Compositor exports video without an audio track. Generate the music in Fuser, download both files, and join them in your video editor. MMAudio and Mirelo SFX can generate sound for an existing clip, but they are built for synchronized sound effects rather than a score. ### How do I make AI music exactly as long as my clip? Use MiniMax Music 3 and set Duration to the clip length. It is a maximum, so check the Actual Duration output. In our test, Duration 10 returned a 10.00-second track for a 10.04-second clip. ### Can Lyria 3.5 make a 10-second track? Not reliably. Lyria 3.5 has no duration setting, only prompt hints. Asking for "a 10-second track" gave us 62.5 seconds, so plan to cut a section in your editor. ### How do I make an instrumental with MiniMax Music 3 in Fuser? The Lyrics field is required, so enter a structure tag such as [instrumental] on its own and write "No vocals. Purely instrumental." in the prompt. Speech-to-text found no words in our result. ### Can a video model use my own music? Wan 3.0 and Seedance 2.5 accept reference audio. In our single test, Wan 3.0 returned a video whose soundtrack was our track, essentially unchanged, while Seedance 2.5 kept the notes but stretched them out by about 12 to 14 percent. Both generate a new shot rather than adding music to an existing clip. ### Do video models with native audio make music? They can. Wan 3.0 is documented as generating background music. Kling describes Kling O3's audio as dialogue, ambience and sound effects, but when our prompt asked for a piano melody its soundtrack contained pitched notes throughout. Either way, the music is the model's, not a track you chose. ## Score the shot next to the shot. Generate the clip and its music side by side, compare scores against the picture, then take the winner to your edit. [Open MiniMax Music 3 in Fuser](https://fuser.studio/models/minimax-music-3) · [Explore all guides](https://fuser.studio/articles) ## More articles - [Lyria 3.5 vs MiniMax Music 3: Same Lyrics, Measured Results](https://fuser.studio/articles/lyria-3-5-vs-minimax-music-3.md) - [How to Make AI Videos Longer: Extend, Chain and Multi-Shot](https://fuser.studio/articles/extend-ai-video-length.md) - [Best AI Video Editing Models: Same Clip, Same Edits](https://fuser.studio/articles/best-ai-video-editing-models.md)