Speech Synthesis on LJSpeech (test)
0.011RTFFastSpeech 2 + HiFiGAN
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FastSpeech 2 + HiFiGANGPU=NVIDIA V100, Batch size=1 sentence2022.05 | 0.011 | — | — | |
| NaturalSpeechGPU=NVIDIA V100, Batch size=1 sentence, Number of parameters=28.7M2022.05 | 0.013 | — | — | |
| VITSGPU=NVIDIA V100, Batch size=1 sentence2022.05 | 0.014 | — | — | |
| Glow-TTS + HiFiGANGPU=NVIDIA V100, Batch size=1 sentence2022.05 | 0.021 | — | — | |
| Grad-TTS (10) + HiFiGANGPU=NVIDIA V100, Batch size=1 sentence, Inference steps=102022.05 | 0.082 | — | — | |
| Grad-TTS (1000) + HiFiGANGPU=NVIDIA V100, Batch size=1 sentence, Inference steps=10002022.05 | 4.12 | — | — | |
| FastSpeechHardware=12 Intel Xeon CPU, 256GB memory, 1 NVIDIA V100 GPU, Batch size=1, Vocoder=WaveGlow, Output type=Mel + WaveGlow2019.05 | — | — | 38.3 | |
| GLA-GradBatch size=100, Hardware=NVIDIA V100 GPU2025.11 | — | — | 32.98 | |
| GLA-Grad++Batch size=100, Hardware=NVIDIA V100 GPU2025.11 | — | — | 37.8 | |
| WaveGradBatch size=100, Hardware=NVIDIA V100 GPU2025.11 | — | — | 42.02 |