F5 TTS

bySWivid (Shanghai AI Lab)

Expressive zero-shot voice cloning and natural speech synthesis powered by flow matching

How F5 TTS works

Produce studio-grade voice clones in three straightforward steps by pairing a short reference recording with structured text.

F5 TTS: Upload reference audio. Real local Fuser canvas capture.

Upload reference audio

Provide a clean, mono 10-to-15-second voice sample with steady pacing and minimal background noise.

F5 TTS: Enter transcript and script. Real local Fuser canvas capture.

Enter transcript and script

Add the reference transcript to bypass ASR, then type your target text with expressive punctuation.

F5 TTS: Synthesize target speech. Real local Fuser canvas capture.

Synthesize target speech

Select F5-TTS or E2-TTS architecture to generate natural speech preserving the original vocal identity.

What F5 TTS is good at

State-of-the-art flow matching and multilingual training deliver organic speech synthesis, accurate timbre cloning, and flexible generation modes.

Zero-Shot Vocal Cloning

Replicate distinct vocal timbres, room acoustic tone, and conversational cadence from a single 10-to-15-second mono reference clip without model fine-tuning.

Cross-Lingual Speech Synthesis

Carry a speaker's unique vocal identity across English, Chinese, French, German, Japanese, and Korean using representations trained on the 100,000-hour Emilia dataset.

Dual Architecture Engine

Switch between the default F5-TTS model for maximum pronunciation robustness and micro-inflection quality, or E2-TTS for rapid conversational prototyping.

Flow-Matching Cadence Control

Harness non-autoregressive flow matching with ConvNeXt V2 blocks to generate natural breathing pauses, micro-inflections, and clean rhythm guided by punctuation.

Made with F5 TTS

Hear how F5 TTS handles varied narrative styles, dramatic line deliveries, and conversational pacing across distinct creative formats.

Documentary nature narration

Cinematic dialogue exchange

Motorsport commercial voiceover

Audiobook narrative reading

What people build with F5 TTS

From serialized audiobooks and localized video to game scratch tracks, creators use F5 TTS to produce expressive, identity-preserving vocal audio.

Audiobook and Narrative Publishing

01

Produce multi-character audiobooks and dramatic narrations with consistent speaker timbre, natural breathing gaps, and organic pacing.

Multilingual Vocal Localization

02

Localize marketing campaigns, educational courses, and corporate presentations into multiple languages while preserving the original speaker's signature voice.

Film and Game Dialogue Prototyping

03

Rapidly draft character dialogue, scratch tracks, and director cuts with cinematic inflections before final voice-talent recording.

Short-Form Social Narration

04

Generate punchy, high-clarity voiceovers for vertical video, product showcases, and tutorials with rapid turnaround.

Virtual Assistants and E-Learning

05

Deploy warm, conversational synthetic speech for interactive learning modules and customer service avatars without robotic artifacts.

F5 TTS specs

Plans and credits
Versions
F5-TTS (Better quality, slower), E2-TTS (Faster, lower quality)
Inputs
Text, Reference Audio (audio), Reference Text (text)
Output
audio
Cost
1 credits per run

Frequently Asked Questions

F5-TTS requires a clean, noise-free mono audio clip between 10 and 15 seconds long with natural speaking pacing. Providing the exact verbatim transcript in the reference text input allows the model to bypass automated speech recognition, accelerating generation speed and preventing pronunciation errors.

F5-TTS offers higher sound fidelity, smoother pronunciation, and superior robustness against skipped words, making it the recommended default choice. E2-TTS generates audio faster with fewer computational steps, but is best reserved for simple sentences and rapid drafting because it is more prone to pronunciation artifacts on complex phrasing.

Yes, F5-TTS supports cross-lingual voice cloning across English, Chinese, German, French, Japanese, and Korean. Because it is trained on the 100,000-hour Emilia dataset, a speaker recorded in English can generate fluent French or Japanese speech while retaining their original vocal timbre and acoustic identity.

Speech distortion or unnatural speed acceleration typically happens when input text chunks exceed 30 seconds of spoken time or when the reference audio contains background noise. Breaking long scripts into smaller segments under 150 characters and using punctuation like commas and periods ensures clean cadence and stable pitch.

F5-TTS was developed by SWivid and researchers at Shanghai AI Lab, first released in October 2024. It utilizes a non-autoregressive Flow Matching framework combined with a Diffusion Transformer backbone and ConvNeXt V2 blocks.

Try F5 TTS on Fuser

Expressive zero-shot voice cloning and natural speech synthesis powered by flow matching