Voice Cloning on SEED-TTS EN (test)
0.99WERFish Audio S2
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Fish Audio S22026.03 | 0.99 | — | — | — | |
| Fish Audio S2Parameters=4B, Open-source=true, Zero-shot=true2026.06 | 0.99 | — | — | — | |
| Fish Audio S12026.03 | 1.07 | — | — | — | |
| Qwen3-TTSParameters=1.7B, Open-source=true, Zero-shot=true2026.06 | 1.23 | 0.717 | — | — | |
| Qwen3-TTS2026.03 | 1.24 | — | — | — | |
| Qwen3-Omni2026.03 | 1.39 | — | — | — | |
| Qwen3-OmniParameters=30B-A3B, Open-source=true, Zero-shot=true2026.06 | 1.39 | — | — | — | |
| LongCat-Audio-DiTParameters=3.5B, Open-source=true, Zero-shot=true2026.06 | 1.5 | 0.786 | — | — | |
| CosyVoice3.5Open-source=false, Zero-shot=true2026.06 | 1.57 | 0.738 | — | — | |
| OmniVoiceParameters=0.8B, Open-source=true, Zero-shot=true2026.06 | 1.6 | 0.741 | — | — | |
| ZipVoiceParameters=0.1B, Open-source=true, Zero-shot=true2026.06 | 1.64 | 0.668 | — | — | |
| MiniMax-SpeechOpen-source=false, Zero-shot=true2026.06 | 1.65 | 0.692 | — | — | |
| DiTARParameters=0.6B, Open-source=false, Zero-shot=true2026.06 | 1.69 | 0.735 | — | — | |
| BareWave (proposed training scheme)Training data=19.4k EN, Params=983.6M, Zero-shot=true2026.06 | 1.75 | — | 0.602 | 3.72 | |
| F5-TTSParams=0.3B, Category=Multi-Stage or NAR Methods, Independent control of timbre and prosody=false2025.12 | 1.83 | 0.647 | — | — | |
| VoxCPM2Parameters=2B, Open-source=true, Zero-shot=true2026.06 | 1.84 | 0.753 | — | — | |
| MOSS-TTSParameters=8B, Open-source=true, Zero-shot=true2026.06 | 1.85 | 0.734 | — | — | |
| VoxCPMParameters=0.6B, Open-source=true, Zero-shot=true2026.06 | 1.85 | 0.729 | — | — | |
| Ground TruthZero-shot=true2026.06 | 1.86 | — | 0.734 | 3.53 | |
| Minimax Speech-022026.03 | 1.9 | — | — | — | |
| OpenAudio-s1-miniParameters=0.5B, Open-source=true, Zero-shot=true2026.06 | 1.94 | 0.55 | — | — | |
| FireRedTTS-22026.03 | 1.95 | — | — | — | |
| FireRedTTS-2Parameters=1.5B, Open-source=true, Zero-shot=true2026.06 | 1.95 | 0.665 | — | — | |
| Spark-TTSParams=0.5B, Category=One-Stage AR Methods, Independent control of timbre and prosody=false2025.12 | 1.98 | 0.584 | — | — | |
| F5-TTSParameters=0.3B, Open-source=true, Zero-shot=true2026.06 | 2 | 0.67 | — | — | |
| CosyVoice 3Parameters=0.5B, Open-source=true, Zero-shot=true2026.06 | 2.02 | 0.718 | — | — | |
| F5-TTS Base (+ Vocos)Training data=19.4k EN, Params=335.8M + 13.5M, Zero-shot=true2026.06 | 2.09 | — | 0.573 | 3.83 | |
| VoxCPM1.5Parameters=0.8B, Open-source=true, Zero-shot=true2026.06 | 2.12 | 0.714 | — | — | |
| CosyVoice 3-1.5BParameters=1.5B2026.03 | 2.21 | — | — | — | |
| CosyVoice 3Parameters=1.5B, Open-source=false, Zero-shot=true2026.06 | 2.22 | 0.72 | — | — | |
| Index-TTS 2Params=1.5B, Category=Multi-Stage or NAR Methods, Independent control of timbre and prosody=true2025.12 | 2.23 | 0.706 | — | — | |
| IndexTTS2Parameters=1.5B, Open-source=true, Zero-shot=true2026.06 | 2.23 | 0.706 | — | — | |
| Seed-TTS2026.03 | 2.25 | — | — | — | |
| Seed-TTSOpen-source=false, Zero-shot=true2026.06 | 2.25 | 0.762 | — | — | |
| BareWave (basic training)Training data=19.4k EN, Params=983.6M, Zero-shot=true2026.06 | 2.34 | — | 0.478 | 3.43 | |
| Simple direct-wave baselineTraining data=19.4k EN, Params=983.6M, Zero-shot=true2026.06 | 2.42 | — | 0.424 | 3.35 | |
| HiggsAudio-v2Parameters=3B, Open-source=true, Zero-shot=true2026.06 | 2.44 | 0.677 | — | — | |
| CosyVoice 2Training data=167K Multi., Params=618M, Zero-shot=true2026.06 | 2.51 | — | 0.659 | 4.15 | |
| VevoParams=N/A, Category=Multi-Stage or NAR Methods, Independent control of timbre and prosody=true2025.12 | 2.53 | 0.664 | — | — | |
| CosyVoice2Params=0.5B, Category=Multi-Stage or NAR Methods, Independent control of timbre and prosody=false2025.12 | 2.57 | 0.652 | — | — | |
| MaskGCTParameters=1B, Open-source=true, Zero-shot=true2026.06 | 2.62 | 0.717 | — | — | |
| Qwen2.5-OmniParameters=7B, Open-source=true, Zero-shot=true2026.06 | 2.72 | 0.632 | — | — | |
| MegaTTS3Parameters=0.5B, Open-source=false, Zero-shot=true2026.06 | 2.79 | 0.771 | — | — | |
| DisCo-SpeechParams=1.5B, Category=One-Stage AR Methods, Independent control of timbre and prosody=true2025.12 | 3.01 | 0.597 | — | — | |
| VibeVoiceParameters=1.5B, Open-source=true, Zero-shot=true2026.06 | 3.04 | 0.689 | — | — | |
| CosyVoice 2Parameters=0.5B, Open-source=true, Zero-shot=true2026.06 | 3.09 | 0.659 | — | — | |
| Spark-TTSParameters=0.5B, Open-source=true, Zero-shot=true2026.06 | 3.14 | 0.573 | — | — | |
| Llasa-1BParams=1B, Category=One-Stage AR Methods, Independent control of timbre and prosody=false2025.12 | 3.22 | 0.572 | — | — | |
| E2-TTS Base (+ Vocos)Training data=19.4k EN, Params=333M + 13.5M, Zero-shot=true2026.06 | 3.5 | — | 0.582 | 3.41 | |
| FireRedTTSParams=N/A, Category=Multi-Stage or NAR Methods, Independent control of timbre and prosody=false2025.12 | 3.82 | 0.46 | — | — | |
| FireRedTTSParameters=0.5B, Open-source=true, Zero-shot=true2026.06 | 3.82 | 0.46 | — | — | |
| FireRedTTSTraining data=248K Multi., Params=~580M, Zero-shot=true2026.06 | 3.82 | — | 0.46 | — | |
| CosyVoiceParameters=0.3B, Open-source=true, Zero-shot=true2026.06 | 4.29 | 0.609 | — | — |