Audio-to-Video Talking Head Generation on Chinese-dominant face video dataset (test)
16.17FIDSonic
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Sonic2026.04 | 16.17 | 106.57 | 1.85 | 11.36 | |
| Talker-T2AV2026.04 | 17.32 | 107.09 | 3.97 | 10.09 | |
| Ditto2026.04 | 17.98 | 187.54 | 1.77 | 11.81 | |
| AniPortrait2026.04 | 23.63 | 336.8 | 1.14 | 12.42 | |
| FLOAT2026.04 | 29.71 | 222.52 | 2.96 | 10.11 | |
| EchoMimic2026.04 | 33.43 | 273.65 | 2.19 | 10.88 |