Talking-head synthesis on Conver-3D YouTube (test)
2.76FDDViBES
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ViBESType=Multi Task, Trained on TFHP=true2025.12 | 2.76 | 8.27 | 12.88 | |
| FaceDiffuserType=Audio → Face, Trained on TFHP=false2025.12 | 2.96 | 9.93 | 14.62 | |
| SelfTalkType=Audio → Face, Trained on TFHP=false2025.12 | 11.56 | 25.45 | 23.44 | |
| ARTalkType=Audio → Face, Trained on TFHP=true2025.12 | 11.6 | 9.04 | 14.42 | |
| CodeTalkerType=Audio → Face, Trained on TFHP=false2025.12 | 17.72 | 10.9 | 16.19 | |
| ScanTalkType=Audio → Face, Trained on TFHP=false2025.12 | 21.02 | 15.81 | 35.08 | |
| MultiTalkType=Audio → Face, Trained on TFHP=false2025.12 | 24.42 | 2.48 | 16.2 | |
| DiffPoseTalkType=Audio → Face, Trained on TFHP=true2025.12 | 28.4 | 8.41 | 12.34 | |
| UniTalkerType=Audio → Face, Trained on TFHP=false2025.12 | 29.31 | 2.07 | 15.3 |