environment-aware text-to-speech on AudioCaps (test)
6.76WERCosyVoice2 + TangoFlux
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CosyVoice2 + TangoFluxMixing separately generated speech and background audio=true2026.05 | 6.76 | — | — | — | 4.01 | 0.452 | |
| ImmersiveTTS#Param.=450M, NFEs=252026.05 | 8.06 | 4.2 | 3.48 | 3.47 | 5.8 | 0.308 | |
| ImmersiveTTS2026.05 | 8.06 | — | — | — | 5.8 | 0.308 | |
| VoiceDiT#Param.=566M, NFEs=2002026.05 | 11.68 | 3.47 | 3.44 | 2.63 | 9.07 | 0.263 | |
| VoiceDiT2026.05 | 11.68 | — | — | — | 9.07 | 0.263 | |
| VoiceLDM#Param.=508M, NFEs=2002026.05 | 16.45 | 3.41 | 3.33 | 2.55 | 8.75 | 0.229 | |
| VoiceLDM2026.05 | 16.45 | — | — | — | 8.75 | 0.229 | |
| Ground Truth2026.05 | 22.29 | — | — | — | — | 0.503 | |
| Reconstructed2026.05 | 22.58 | 4.08 | 4.16 | 3.49 | — | 0.488 | |
| AudioLDM2 (Speech)2026.05 | 35.06 | — | — | — | 29.59 | 0.048 | |
| AudioLDM2 (Speech) + AudioLDM2 (Audio)Mixing separately generated speech and background audio=true2026.05 | 41.33 | — | — | — | 5.36 | 0.365 |