Text-to-Speech on ESD (test)
4.47MOSReference
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| ReferenceEvaluation Mode=Parallel2022.05 | 4.47 | — | — | — | — | |
| Reference(voc.)Evaluation Mode=Parallel, Vocoder=HiFi-GAN (V1)2022.05 | 4.4 | 4.47 | 0.99 | 0.07 | — | |
| Recording2026.06 | 4.21 | — | — | — | — | |
| EmoShift2026.01 | 4.14 | — | — | — | 3.96 | |
| GenerSpeechEvaluation Mode=Parallel, Batch size=1, Vocoder=HiFi-GAN (V1)2022.05 | 4.11 | 4.2 | 0.97 | 0.26 | — | |
| CosyVoicemode=Zero-Shot2026.01 | 4.07 | — | — | — | 3.67 | |
| dynamic prosody predictionTraining dataset=50k hours, alpha=0.52026.06 | 4.07 | — | — | — | — | |
| Expressive FS2Evaluation Mode=Parallel, Batch size=1, Vocoder=HiFi-GAN (V1)2022.05 | 4.04 | 3.93 | 0.93 | 0.41 | — | |
| Meta-StyleSpeechEvaluation Mode=Parallel, Batch size=1, Vocoder=HiFi-GAN (V1)2022.05 | 4.02 | 3.97 | 0.86 | 0.41 | — | |
| CosyVoice(50k)Training dataset=50k hours2026.06 | 4.01 | — | — | — | — | |
| CoTTraining dataset=50k hours2026.06 | 4 | — | — | — | — | |
| CosyVoice-SFTmode=Supervised Fine-Tuned2026.01 | 3.93 | — | — | — | 3.79 | |
| MellotronEvaluation Mode=Parallel, Batch size=1, Vocoder=HiFi-GAN (V1)2022.05 | 3.92 | 4.01 | 0.8 | 0.27 | — | |
| FG-TransformerTTSEvaluation Mode=Parallel, Batch size=1, Vocoder=HiFi-GAN (V1)2022.05 | 3.9 | 3.94 | 0.67 | 0.43 | — | |
| StylerEvaluation Mode=Parallel, Batch size=1, Vocoder=HiFi-GAN (V1)2022.05 | 3.76 | 4.05 | 0.68 | 0.39 | — |