Text-to-speech on SeedTTS en (eval)
1.1WERAudioCALM
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AudioCALMBaseline Type=Unified2026.06 | 1.1 | 67.2 | 3.95 | |
| Ming-omni-TTSBaseline Type=Unified2026.06 | 1.3 | 63.3 | 3.8 | |
| CosyVoice 3.0Baseline Type=Modality-specific2026.06 | 1.5 | 69.5 | 3.88 | |
| F5-TTSBaseline Type=Modality-specific2026.06 | 1.8 | 64.8 | 3.78 | |
| UniMoE-AudioBaseline Type=Unified2026.06 | 1.9 | 57.3 | 3.72 | |
| UniFlow-AudioBaseline Type=Unified2026.06 | 5.8 | 57.3 | 3.45 | |
| UniAudioBaseline Type=Unified2026.06 | 11.3 | 36.3 | 3.22 |