Description-based speech generation on Accent+ (test)
0.596JointCLAPAUDIOBOX
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| AUDIOBOX2023.12 | 0.596 | 2.6 | — | — | |
| ground truth2023.12 | 0.561 | 13.5 | — | — | |
| VoiceLDM2023.12 | 0.235 | 4.4 | — | — | |
| AudioLDM2-SP2023.12 | 0.11 | 23.9 | — | — |
| Method | Links | ||||
|---|---|---|---|---|---|
| AUDIOBOX2023.12 | 0.596 | 2.6 | — | — | |
| ground truth2023.12 | 0.561 | 13.5 | — | — | |
| VoiceLDM2023.12 | 0.235 | 4.4 | — | — | |
| AudioLDM2-SP2023.12 | 0.11 | 23.9 | — | — |