Video-to-Speech Synthesis on V2C-Animation (test)
0MCDGround Truth
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Ground Truth2021.11 | 0 | 0 | 0 | 90.62 | 84.38 | — | — | — | — | |
| Ground Truth2022.12 | 0 | 0 | 0 | 90.62 | 84.38 | 4.61 | 4.74 | 6.734 | 7.813 | |
| V2C-Net2021.11 | 11.79 | 10.09 | 10.05 | 62.5 | 56.25 | — | — | — | — | |
| FastSpeech22021.11 | 12.08 | 10.29 | 10.31 | 59.38 | 53.13 | — | — | — | — | |
| Ours2022.12 | 15.66 | 12.29 | 13.48 | 37.75 | 61.46 | 4.03 | 3.89 | 8.036 | 5.608 | |
| SV2TTS*video (emotion) embedding=true2021.11 | 17.41 | 11.16 | 15.92 | 38.21 | 41.24 | — | — | — | — | |
| SV2TTSVideo (emotion) embedding=true2022.12 | 19.38 | 12.73 | 34.51 | 35.18 | 42.05 | 2.07 | 2.15 | 12.617 | 3.349 | |
| TacotronVideo (emotion) embedding=true2022.12 | 19.79 | 18.73 | 42.15 | 32.49 | 39.68 | 2.12 | 2.06 | 13.475 | 2.938 | |
| V2C-Net2022.12 | 20.61 | 14.23 | 19.15 | 26.84 | 48.41 | 3.19 | 3.06 | 11.784 | 3.026 | |
| FastSpeech2Video (emotion) embedding=true2022.12 | 20.66 | 14.59 | 20.79 | 23.44 | 46.9 | 3.08 | 2.89 | 12.113 | 2.604 | |
| FastSpeech2Video (emotion) embedding=false2022.12 | 20.78 | 14.39 | 19.41 | 21.72 | 46.82 | 2.79 | 2.63 | 12.261 | 2.958 | |
| SV2TTSvideo (emotion) embedding=false2021.11 | 21.08 | 12.87 | 49.56 | 33.62 | 37.19 | — | — | — | — | |
| SV2TTSVideo (emotion) embedding=false2022.12 | 21.08 | 12.87 | 49.56 | 33.62 | 37.19 | 2.03 | 1.92 | 13.733 | 2.725 |