ElevenLabs Sound Effects Prompt Guide: Describe, Time and Sync Your SFX

A tested guide to ElevenLabs Sound Effects v2: how to describe a sound, when to set a duration, what prompt influence changes, what happens with loops, and when a video-to-audio model like Mirelo or MMAudio is the better tool.

FuserUpdated
Fuser canvas: a sound prompt feeds ElevenLabs SFX, and a silent kiln clip plus the same prompt feed Mirelo SFX and MMAudio, each showing its real waveform.

All guides · ElevenLabs SFX in Fuser

Quick answer: prompt ElevenLabs Sound Effects with a short, concrete description of what makes the sound, what it is made of, and where it happens: "heavy wooden door creaking open" beats "door". For anything with more than one event, write the events in order with "then" and set a duration long enough to hold them, because the automatic length can be too short. Keep Prompt Influence near the 0.3 default for variety and raise it when you need the model to stay literal. If you already have the video, a video-to-audio model such as Mirelo SFX or MMAudio will time the sound to the picture for you.

How to describe a sound

ElevenLabs' own guidance splits prompts into two kinds (ElevenLabs sound effects docs):

  • Simple effects: one clear, concise description, such as "glass shattering on concrete" or "thunder rumbling in the distance".

  • Sequences: the events in the order they happen, such as "footsteps on gravel, then a metallic door opens".

In practice a good one-line prompt answers four questions:

  1. Source. What makes the sound: a kiln door, a car door, a bike bell.

  2. Material and weight. Iron hinges, gravel, thin glass, a heavy door. This is what separates a creak from a clack.

  3. Action and order. Opens, slams, rolls, then. One verb per event.

  4. Space. A quiet studio, a tiled bathroom, an open field, or a distance cue like ElevenLabs' own "thunder rumbling in the distance". Name the place once rather than describing a whole scene.

The same docs list sound-design terms to use in prompts, and they are worth using because they are short and exact: impact, whoosh, ambience, one-shot, loop, stem, braam, glitch and drone. The model also takes musical elements, for example a drum loop with a tempo in BPM or a synth pad, but for full tracks use a music model such as MiniMax Music.

The fal endpoint Fuser calls accepts up to 450 characters of prompt (fal model page). That is plenty; if you are near the limit, you are probably describing a scene rather than a sound.

Duration: when to set it

The ElevenLabs SFX node in Fuser has an optional Duration field from 0.5 to 22 seconds. Leave it empty and the model chooses a length from the prompt. We ran the same kinds of prompts both ways:

Real outputs, drawn at their actual level. An empty Duration let the model pick 6 s for “door” but only 1 s for a two-event prompt. Generated 28 September 2026.
  • "door", no duration: a 6.0-second file with one hard hit at about 4 seconds, a faint tick near the start, and near-silence otherwise. In this run the vague prompt got a long, mostly empty file.

  • The kiln prompt ("heavy kiln door creaks open on iron hinges, then a low roar of heat from inside, quiet pottery studio"), no duration: 1.0 second. That is too short to hold a creak followed by a roar.

  • The same kiln prompt at 5 seconds: about 3 seconds of texture that then fades out, and quiet: it peaks around −21 dBFS. Plan to raise the gain in your editor.

  • "Three slow footsteps on gravel, a pause, then a car door slams shut", no duration: 2.0 seconds, with two hits in the first half-second, a gap, and the loudest hit at 1.4 seconds. The order came through, but we asked for three footsteps and the waveform shows two, and the pacing is tighter than "slow" suggests.

  • "Soft UI confirmation click, single short one-shot, clean, no reverb" at 0.5 seconds: a 0.48-second file that has decayed by about 0.2 seconds, ready to drop onto a button.

The rule that falls out: leave Duration empty for single hits, set it for sequences and for anything that has to fill a shot. For ambience we set the maximum, 22 seconds, on "Ambience: steady rain on a tin roof, distant thunder rolls twice, no music": the level stayed even for all 22 seconds and the two thunder rolls did not show up as separate peaks, so listen for timed events rather than assuming they landed.

Prompt Influence

Prompt Influence runs from 0 to 1 and defaults to 0.3. ElevenLabs describes high values as a more literal reading of the prompt and low values as a more creative one with more variation (ElevenLabs sound effects docs). The node has no seed, so the practical workflow is to run the node a few times at the default, keep the take you like, and only push the slider up when the model keeps drifting away from what you wrote.

Loops

ElevenLabs' API has a loop option that makes a sound effect repeat without a perceptible start or end (ElevenLabs API reference). Fuser's ElevenLabs SFX node does not expose it. In Fuser, generate the longest take you need (up to 22 seconds) and loop it in your editor, checking the join by ear.

Output format

The node's Output Format menu offers MP3 at 64, 128 (default) or 192 kbps, Opus at 128 kbps, and uncompressed PCM at 44.1 kHz. Use PCM when the sound goes into further editing, MP3 128 for quick previews.

Text-to-SFX vs video-to-audio

ElevenLabs SFX only reads text, so it does not know when something happens on screen. For a clip you already have, Fuser has two nodes that watch the video. We sent the same silent 5-second kiln clip and the same prompt to both, and ran the prompt through ElevenLabs at 5 seconds for comparison:

One silent 5-second clip and one prompt. The two video-to-audio models got loud when the door moved; ElevenLabs had no picture to follow. Generated 28 September 2026.
  • Mirelo SFX 1.6 takes a video and an optional prompt and returns the video with a new effects track. In Fuser it runs 1 to 60 seconds (default 10, so set it to your clip length), with 1 to 4 variations per run and a seed. fal notes that durations over 10 seconds use sliding-window generation (fal model page). In our run it stayed quiet for the first 1.25 seconds and rose as the door swung.

  • MMAudio V2 does the same job from 1 to 30 seconds (default 5), adds a negative prompt (default "noisy, low quality, low volume"), and switches to its text-to-audio endpoint when no video is connected (fal model page). With a video it returns the video with sound. Its authors note that its default output and training duration is 8 seconds (the Fuser node defaults to 5), that it sometimes produces unintelligible speech-like sounds, and that it struggles with unfamiliar concepts (MMAudio on GitHub). In our run it also rose with the door, and its level kept moving through the rest of the clip.

Both video-to-audio outputs came back far louder than the text-only take (both peaked within about 1.5 dB of full scale, against about −21 dBFS for the text-only take). Use ElevenLabs when you need a clean, specific sound to place yourself: UI sounds, a signature impact, a whoosh for a cut, a bed of ambience. Use Mirelo or MMAudio when the timing has to match what the camera shows. For a wider roundup, see the best AI music and sound effect generators.

Build it on the canvas

In Fuser, keep the prompt in its own text node and wire it into the ElevenLabs SFX node, and into Mirelo or MMAudio when you also have the clip, as in the canvas at the top of this page. Generate the picture upstream with an image-to-video model, then use Download media on the node to take the audio or the finished video into your editor. For voice lines, add ElevenLabs TTS on the same canvas; for the whole chain from still to sound, see how to chain AI models.

Tested 28 September 2026 on the fal endpoints Fuser uses.

ElevenLabs SFX prompt cheat sheet.

What to write for common sound types, and the settings that go with them in Fuser.

SoundWriteExample
Prompt patterns
Impact

Object, material, surface, space.

Ceramic bowl dropped on a stone floor, small room

Whoosh

Speed and texture of the movement.

Fast airy whoosh past the camera

Ambience

Place, weather, a few steady details. Set a long duration.

Steady rain on a tin roof, 22 s

Sequence

Events in order, joined with then. Set a duration.

Footsteps on gravel, then a car door slams

UI one-shot

Short, clean, dry, one-shot.

Soft UI confirmation click, no reverb, 0.5 s

Settings in Fuser
Duration

Empty for single hits; set it for sequences and beds.

0.5 to 22 seconds

Prompt Influence

Keep 0.3 for variety; raise it if takes drift.

0 to 1, default 0.3

Output Format

PCM for editing, MP3 128 for previews.

MP3 64/128/192, Opus 128, PCM 44.1 kHz

Loop

Not exposed on the node; loop in your editor.

API-only option

Questions, answered.

Describe the source, its material, the action and the space in one line, for example "heavy wooden door creaking open in a stone hallway". For several events, list them in order with "then". Sound-design terms such as impact, whoosh, ambience and one-shot help.

In Fuser, 0.5 to 22 seconds. Leave Duration empty and the model picks a length; ElevenLabs' own API allows up to 30 seconds.

ElevenLabs' API has a loop option, but Fuser's ElevenLabs SFX node does not expose it. Generate a longer take and loop it in your editor, then check the join by ear.

It sets how literally the model follows your text, from 0 to 1. The default is 0.3. Higher values follow the prompt more closely with less variation; lower values give more varied takes.

If you already have the clip, use a video-to-audio model such as MMAudio or Mirelo SFX so the sound follows the action. In our test both got loud when the on-screen door opened. Use ElevenLabs SFX for specific sounds you will place yourself.

Give every shot its sound.

Prompt effects, score silent clips and keep the whole chain on one canvas.

All articles