# MMAudio Canonical page: https://fuser.studio/models/mmaudio MMAudio in Fuser generates audio synced to a video, 1-30 seconds long, from the clip and an optional text prompt. Add sound to silent generated video. ## Specs | Spec | Value | | --- | --- | | Inputs | Prompt (text), Video, Negative Prompt (text) | | Output | audio, video | | Duration | 1–30 seconds | | Cost | 1–41 credits per run | Generate studio-grade 44.1kHz Foley and physical soundscapes synchronized to silent video frames or synthesized from text Creator: WaveSpeed AI Also known as: MMAudio, MMAudio V2 [Generate Audio](https://app.fuser.studio) ## How MMAudio works Turn silent AI video generations or descriptive prompts into finished, frame-locked audio tracks in seconds. ### Input Video or Prompt Provide a silent video clip to synchronize Foley against motion cues, or enter a text prompt describing the desired acoustic textures. ### Adjust Acoustic Parameters Set duration between 1 and 30 seconds, select sampling steps, and tune guidance scale to balance prompt fidelity and visual timing. ### Export Synchronized Audio Receive a full-fidelity 44.1kHz audio track or multiplexed video file with transient peaks aligned to on-screen motion. ## What MMAudio is good at A 157M flow matching audio synthesis engine built for frame-accurate Foley and tactile physical soundscapes. ### Frame-Accurate Action Foley Extract visual features at 24 fps to lock acoustic hits and transient peaks precisely onto visual impact frames in silent video clips. ### Immersive Environmental Ambience Generate high-density, multi-layered acoustic beds with natural environmental reverberation, from dripping industrial tunnels to storm-swept coastal winds. ### Standalone Text-to-Sound Synthesis Synthesize crisp physical friction and impact textures without needing video input, ideal for building bespoke SFX libraries directly from text descriptions. ### High-Fidelity 44.1kHz Acoustic Output Deliver clean 44.1kHz audio waveforms that preserve subtle high-frequency details, transient snap, and organic physical decays. ## Made with MMAudio Listen to real-world acoustic textures, impact synchronization, and ambient beds generated across diverse production formats. ### Cinematic film impact Foley ### Documentary nature soundscape ### Industrial machine sound design ### Culinary e-commerce Foley bed ### Sci-fi short film mechanical asset ## What people build with MMAudio From film post-production to game asset design, creators rely on MMAudio to add physical acoustic weight to silent visuals. ### Cinematic Film and Video Post-Production Eliminate manual Foley spotting by generating frame-synchronized footsteps, body impacts, and atmospheric beds directly from silent video renders. ### Game Sound Design and Asset Generation Create tactile material sounds, UI feedback hits, and realistic object collision audio for interactive environments without field recording sessions. ### Social Media and Short-Form Commercials Pair social video reels and product showcases with hyper-tactile cooking, unboxing, or mechanical textures that heighten viewer engagement. ### Documentary and Archival Media Soundscapes Construct convincing outdoor and indoor soundscapes with natural acoustic reverberation for background beds in non-fiction media. ### E-Commerce and Product Demonstration Video Synthesize crisp product interaction sounds—like clicking switches, pouring liquids, and opening containers—to give e-commerce video tactile weight. ## Frequently Asked Questions ### What is the difference between video-to-audio and text-to-audio modes? The video-synchronized variant analyzes video frames at 24 fps to align acoustic energy peaks directly with on-screen visual motion, while the text-to-audio variant generates unprompted standalone sound effects and ambient soundscapes purely from your text description without requiring video footage. ### Which MMAudio mode should I use for my project? Use video-synchronized mode whenever you have rendered silent footage and need exact frame-accurate hits like footsteps or collisions; use text-to-audio mode when building standalone sound libraries, Foley spot effects, or looping room-tone beds. ### Is MMAudio the same as MMAudio V2? Yes, MMAudio V2 is the 157-million-parameter compact flow matching architecture developed by researchers at UIUC and Sony AI, built specifically to synthesize 44.1kHz Foley and soundscapes synchronized to video. ### What duration range does the model support? MMAudio generates audio tracks from 1 to 30 seconds in duration. It was natively trained on 8-second clips, making 5 to 12 seconds the ideal range for tight visual-to-audio synchronization without timing drift. ### Can MMAudio generate dialogue or background music? MMAudio is engineered for physical Foley and environmental ambience, making it weak at generating coherent human speech, singing vocals, and structured musical scores. For narrative speech or melodic soundtracks, combine MMAudio with dedicated voice or music models. ## Try MMAudio on Fuser Generate studio-grade 44.1kHz Foley and physical soundscapes synchronized to silent video frames or synthesized from text [Generate Audio](https://app.fuser.studio) ## Recipes using MMAudio - [Video Sound Generator](https://fuser.studio/recipes/video-sound-generator.md): Provide a source video and sound description to create an ambient soundtrack and a version of the video with sound. ## Used for - [Film & animation](https://fuser.studio/solutions/industries/film-animation.md): Develop film treatments, character and environment concepts, storyboards and motion studies. Connect image, video and 3D exploration in Fuser. - [Video & storyboards](https://fuser.studio/solutions/use-cases/video-storyboards.md): Develop storyboard frames, shot variations and video studies from your references. Compose clear boards and explore movement in one Fuser project. ## Related models - [MiniMax Music 3](https://fuser.studio/models/minimax-music-3.md) - [Lyria 3.5](https://fuser.studio/models/lyria-3-5.md) - [ElevenLabs TTS](https://fuser.studio/models/elevenlabs-tts.md) - [F5 TTS](https://fuser.studio/models/f5-tts.md) - [ElevenLabs SFX](https://fuser.studio/models/elevenlabs-sfx.md) - [Minimax Music](https://fuser.studio/models/minimax-music.md) Node reference: [MMAudio docs](https://docs.fuser.studio/docs/nodes/audio/mmaudio.md) ## More articles - [AI Music for Video: Score a Generated Clip with Lyria 3.5 or MiniMax Music 3](https://fuser.studio/articles/ai-music-for-video.md) - [How to Edit a Video With Text Prompts: Kling O3 Edit, Gemini Omni, Grok and Luma Tested](https://fuser.studio/articles/edit-video-with-text-prompts.md) - [Wan 3.0 Guide: 30-Second Video, References and Prime](https://fuser.studio/articles/wan-3-0-guide.md)