Zero-shot Text-to-Speech on LibriSpeech-PC clean (test)
1.78WERBigVGAN
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| BigVGANParams (M)=112.4, Mel-type=librosa2025.12 | 1.78 | 0.67 | 4.05 | |
| PeriodWave-TurboParams (M)=70.2, Mel-type=librosa, Inference steps=42025.12 | 1.8 | 0.67 | 4.17 | |
| BigVGAN-v2*Params (M)=112.4, Mel-type=librosa, trained_on_large_scale_dataset=true2025.12 | 1.82 | 0.66 | 3.7 | |
| RFWaveParams (M)=18.1, Mel-type=torchaudio, Inference steps=102025.12 | 1.87 | 0.66 | 3.63 | |
| Flow2GANParams (M)=78.9, Mel-type=torchaudio, Inference steps=4, Noise during GAN fine-tuning=true2025.12 | 1.87 | 0.67 | 4.26 | |
| Flow2GANParams (M)=78.9, Mel-type=torchaudio, Inference steps=1, Noise during GAN fine-tuning=true2025.12 | 1.88 | 0.67 | 3.83 | |
| Flow2GANParams (M)=78.9, Mel-type=torchaudio, Inference steps=2, Noise during GAN fine-tuning=true2025.12 | 1.88 | 0.67 | 4.21 | |
| Flow2GANParams (M)=78.9, Mel-type=torchaudio, Inference steps=4, Noise during GAN fine-tuning=false2025.12 | 1.89 | 0.67 | 4.18 | |
| VocosParams (M)=13.5, Mel-type=torchaudio2025.12 | 1.91 | 0.65 | 3.9 | |
| Flow2GANParams (M)=78.9, Mel-type=torchaudio, Inference steps=2, Noise during GAN fine-tuning=false2025.12 | 1.91 | 0.67 | 4.11 | |
| Flow2GANParams (M)=78.9, Mel-type=torchaudio, Inference steps=1, Noise during GAN fine-tuning=false2025.12 | 1.94 | 0.67 | 3.75 | |
| WaveFMParams (M)=19.5, Mel-type=torchaudio, Inference steps=12025.12 | 2.01 | 0.65 | 3.15 |