Text-to-Speech on LibriSpeech 40 samples subset (test-clean)
1.94WERGround Truth
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Ground TruthSubset=40 samples subset2024.10 | 1.94 | 0.68 | — | |
| NaturalSpeech 3#Param.=500M, #Data=60K EN2024.10 | 1.94 | 0.67 | 0.296 | |
| Voicebox#Param.=330M, #Data=60K EN2024.10 | 2.03 | 0.64 | 0.64 | |
| MaskGCT#Param.=1048M, #Data=100K Multi.2024.10 | 2.634 | 0.687 | — |