Video-to-Speech synthesis on V2C-Animation
79Sim-OGround Truth
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Ground Truth2024.02 | 79 | — | |
| StyleDubberVisual=true2024.02 | 25 | 0.34 | |
| StyleSpeechVisual=false2024.02 | 14 | 0.23 | |
| StyleSpeechVisual=true2024.02 | 14 | 0.23 | |
| Zero-shot TTSVisual=true2024.02 | 13 | 0.22 | |
| Zero-shot TTSVisual=false2024.02 | 12 | 0.21 | |
| HPMDubbingVisual=true2024.02 | 11 | 0.19 | |
| Fastspeech2Visual=false2024.02 | 10 | 0.19 | |
| Fastspeech2Visual=true2024.02 | 10 | 0.18 | |
| Face-TTSVisual=true2024.02 | 9 | 0.12 | |
| V2C-NetVisual=true2024.02 | 8 | 0.15 |