Speech Synthesis on Manchu Speech Dataset (test)
4.68MOSGround Truth
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Ground TruthModel=GT, Description=Human-recorded speech2025.12 | 4.68 | 4.4 | 14.2 | 11.6 | 90.5 | 3.89 | |
| ManchuTTSModel=Ours2025.12 | 4.52 | 5.83 | 18.7 | 12.4 | 84.7 | 3.21 | |
| CBVCModel=CBVC2025.12 | 4.28 | 5.91 | 19.5 | 14.1 | 81.8 | 3.15 | |
| F5-TTSModel=F52025.12 | 4.12 | 5.99 | 20.3 | 15.6 | 79.4 | 3.08 | |
| VITSModel=VITS2025.12 | 3.95 | 6.18 | 21.7 | 17.8 | 76.9 | 3.02 | |
| Glow-TTSModel=Glow2025.12 | 3.89 | 6.34 | 22.4 | 18.9 | 75.6 | 2.97 | |
| FastSpeech 2Model=FS22025.12 | 3.45 | 7.15 | 29.8 | 24.1 | 69.8 | 2.68 | |
| Tacotron 2Model=T22025.12 | 3.21 | 7.92 | 34.2 | 28.7 | 63.1 | 2.45 |