# LatentSync Canonical page: https://fuser.studio/models/latentsync LatentSync in Fuser syncs the lips in a video to a new audio track. Pair it with text-to-speech nodes to dub or revoice footage on one canvas. ## Specs | Spec | Value | | --- | --- | | Inputs | Video, Audio | | Output | video | | Cost | 689 credits per run | Synchronize high-resolution dialogue video to any vocal track with lifelike skin texture and sharp dental fidelity Creator: ByteDance [Sync dialogue video](https://app.fuser.studio) ## How LatentSync works Transform existing footage into naturally synchronized dialogue in three simple stages. ### Upload talking-head video Provide a high-resolution, front-facing video clip with clear lighting and an unobstructed mouth. ### Attach clean vocal audio Supply a clean vocal track or voiceover recording with minimal background noise. ### Tune guidance and generate Adjust the guidance scale between 1.0 and 2.0, select a loop mode, and generate synced video. ## What LatentSync is good at Built on latent diffusion and audio-visual cross-attention, LatentSync preserves natural facial anatomy and skin texture across every frame. ### Latent-Space Facial Synthesis Operates directly inside latent diffusion space to preserve microscopic skin pores, lip texture, and dental geometry without the smearing or blurry mouth boundaries typical of legacy GAN models. ### Temporal Motion Coherence Employs Temporal Representation Alignment to bind speech phonemes to facial motion, eliminating temporal flickering and keeping jaw movement stable across consecutive frames. ### Adjustable Guidance Scale Fine-tune vocal adherence between 1.0 and 2.0, balancing subtle natural lip motion at standard speech tempos against intense articulation for fast dialogue. ### Continuous Video Looping Configure pingpong or standard loop modes to seamlessly cycle talking-head video clips when the input audio duration exceeds the video footage length. ## Made with LatentSync Explore high-fidelity lip-sync examples spanning commercial campaigns, cinematic dubbing, and editorial portraits. ### Editorial spoken-word performance with high-contrast color grading ### Cinematic music-video portrait with atmospheric lighting ### Vertical social-media creator instructional monologue ### E-commerce product presentation with precise consonant sync ## What people build with LatentSync Discover how localization studios, filmmakers, and digital creators dub dialogue without sacrificing visual realism. ### Multilingual Video Localization Dub commercial campaigns into new languages while maintaining exact skin texture and convincing lip movement for global markets. ### Cinematic Dialogue Replacement Replace dialogue in film and television scenes without expensive reshoots, preserving actor facial micro-expressions and high-frequency details. ### Virtual Avatar Broadcasting Drive photorealistic digital human avatars and automated presenters with dynamic voice recordings for daily content streams. ### High-Definition Commercial Dubbing Re-voice high-definition beauty and skincare advertisements where realistic lip contours and sharp teeth rendering cannot look synthetic. ### Educational Lecture Translation Translate academic and training video courses into diverse languages, keeping instructors naturally synchronized with localized lectures. ## Frequently Asked Questions ### What makes LatentSync different from GAN-based lip-sync models? LatentSync performs lip synchronization directly in the latent space of Stable Diffusion rather than relying on pixel-space GANs. This architectural approach eliminates blurry mouth patches and artificial skin rings, preserving fine skin pores, lip creases, and individual dental structures for seamless realism. ### What video and audio inputs produce the highest quality results? Use front-facing source clips at 512x512 resolution or higher with consistent facial lighting and unobstructed mouth views. Pair with isolated mono voice recordings free from heavy background noise or music. Pre-trimming video clips to 5 to 15 seconds prevents temporal drift and maximizes synchronization precision. ### How does the guidance scale setting affect the output? The guidance scale ranges from 1.0 to 2.0 and determines how strictly the diffusion model tracks spoken phonemes. A scale of 1.5 offers the best balance between sharp articulation and facial stability. Lower values toward 1.0 reduce edge tension, while values up to 2.0 enforce rapid consonant articulation. ### When should I choose pingpong loop mode over standard loop mode? Choose pingpong mode for stationary talking-head clips where reversing frames creates an imperceptible back-and-forth flow. Select loop mode when the source clip contains directional head motion or background movements that require continuous one-way visual momentum. ### What types of content should I avoid using with LatentSync? Avoid stylized 2D cartoons, anime characters, and clips with extreme side-profile angles. LatentSync also struggles when objects such as hands, microphones, or cups occlude the mouth, which interferes with face-tracking bounding boxes. ### Who created LatentSync and when was it released? LatentSync was developed by ByteDance and introduced in December 2024 as an open-source latent diffusion framework for audio-conditioned lip-sync animation. ## Try LatentSync on Fuser Synchronize high-resolution dialogue video to any vocal track with lifelike skin texture and sharp dental fidelity [Sync dialogue video](https://app.fuser.studio) ## Related models - [MiniMax H3 Recast](https://fuser.studio/models/minimax-h3-recast.md) - [Kling O3 Edit](https://fuser.studio/models/kling-o3-edit.md) - [Kling O3](https://fuser.studio/models/kling-o3.md) - [Kling 3.0 Video](https://fuser.studio/models/kling-3-0-video.md) - [Seedance 2](https://fuser.studio/models/seedance-2.md) - [MiniMax H3](https://fuser.studio/models/minimax-h3.md) Node reference: [LatentSync docs](https://docs.fuser.studio/docs/nodes/video/latentsync.md) ## More articles - [MiniMax H3 Recast Guide: Replace the People in a Video From Photos](https://fuser.studio/articles/minimax-h3-recast-guide.md) - [MiniMax Music 3 Prompt Guide: Structured Captions, Section Tags and Song Length](https://fuser.studio/articles/minimax-music-3-prompt-guide.md) - [MiniMax Music 3 vs Music-01: Same Lyrics, Measured Side by Side](https://fuser.studio/articles/minimax-music-3-vs-music-01.md)