Diverse Speech Generation on LibriSpeech (test-other)
3.1WERVoicebox (VB-En)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Voicebox (VB-En)Input requirements=text-only, Guidance strength (α)=0, Duration model=regression2023.06 | 3.1 | 155.7 | |
| Ground truth2023.06 | 4.3 | 171.1 | |
| VITS-LJInput requirements=text-only2023.06 | 5.6 | 344.2 | |
| Voicebox (VB-En)Input requirements=text-only, Guidance strength (α)=0, Duration model=flow matching (FM), Duration guidance strength (α_dur)=02023.06 | 5.6 | 159.8 | |
| YourTTSInput requirements=require additional input, Reference audio=LS train2023.06 | 9 | 277.9 | |
| VITS-VCTKInput requirements=require additional input2023.06 | 10.6 | 306.6 | |
| A3TInput requirements=text-only2023.06 | 37.9 | 373 |