Text-to-Speech on LibriSpeech (test-clean)
0.26Log F0 RMSE (Avg)F5-TTS
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| F5-TTSGenerative Paradigm=NAR (flow matching), Zero-shot evaluation=true2025.09 | 0.26 | 2.98 | 3.48 | 2.19 | 79.59 | 1.5 | |
| CosyVoiceGenerative Paradigm=AR (next-token prediction), Zero-shot evaluation=true2025.09 | 0.27 | 3.46 | 4.08 | 4.41 | 120.59 | 4.59 | |
| E2 TTSGenerative Paradigm=NAR (flow matching), Zero-shot evaluation=true2025.09 | 0.27 | 3.26 | 3.4 | 1.87 | 84.91 | 2.11 | |
| MaskGCTGenerative Paradigm=NAR (masked generative modeling), Zero-shot evaluation=true2025.09 | 0.28 | 4.04 | 4.76 | 6.51 | 139.75 | 5.61 | |
| ZipVoiceGenerative Paradigm=NAR (flow matching), Zero-shot evaluation=true2025.09 | 0.29 | 4.55 | 3.91 | 3.85 | 114.52 | 3.93 | |
| CosyVoice 2Generative Paradigm=AR (next-token prediction), Zero-shot evaluation=true2025.09 | 0.3 | 4.56 | 4.07 | 4.47 | 134.34 | 5.38 | |
| XTTS-v2Generative Paradigm=AR (next-token prediction), Zero-shot evaluation=true2025.09 | 0.31 | 5.14 | 4.18 | 4.7 | 127.84 | 4.89 |