Environment-aware Text-to-Speech on Seed-TTS en and AudioCaps augmented (test)
3.59WERReconstructed
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Reconstructed#Param.=-, NFEs=-2026.05 | 3.59 | 4.02 | 3.95 | 3.41 | — | 0.291 | |
| ImmersiveTTS#Param.=450M, NFEs=252026.05 | 4.48 | 4.18 | 3.32 | 3.23 | 3.92 | 0.207 | |
| VoiceDiT#Param.=566M, NFEs=2002026.05 | 7.08 | 3.45 | 3.38 | 3.12 | 5.37 | 0.134 | |
| Ground Truth (Augmented)#Param.=-, NFEs=-2026.05 | 7.86 | — | — | — | — | 0.317 | |
| VoiceLDM#Param.=508M, NFEs=2002026.05 | 11.2 | 3.32 | 3.24 | 2.91 | 6.98 | 0.118 |