Video-to-Speech Synthesis on GRID (test)
0.87Sim-OGround Truth
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Ground Truth2024.02 | 0.87 | — | |
| StyleDubberVisual=true2024.02 | 0.75 | 0.8 | |
| StyleSpeechVisual=false2024.02 | 0.74 | 0.79 | |
| StyleSpeech*Visual=true2024.02 | 0.74 | 0.79 | |
| Zero-shot TTSVisual=false2024.02 | 0.7 | 0.75 | |
| Zero-shot TTS*Visual=true2024.02 | 0.69 | 0.74 | |
| Fastspeech2*Visual=true2024.02 | 0.48 | 0.52 | |
| HPMDubbingVisual=true2024.02 | 0.46 | 0.56 | |
| V2C-NetVisual=true2024.02 | 0.43 | 0.55 | |
| Face-TTSVisual=true2024.02 | 0.42 | 0.51 | |
| Fastspeech2Visual=false2024.02 | 0.38 | 0.42 |