Speech-driven 3D Facial Animation on 3D Face-to-Face Interaction Dataset
10.43Facial Dynamics Distance (FD)Ours
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ours2026.03 | 10.43 | 18.24 | 4.03 | 2.09 | 3.5 | 0.98 | 0.35 | 7.99 | 2.29 | 2.28 | 2.48 | 1.97 | 2.45 | |
| Ours (Single)training=trained on single-person data2026.03 | 19.58 | 29.03 | 6.32 | — | 6.74 | 1.23 | 1.14 | 6.86 | 5.23 | 2.23 | 1.4 | 1.61 | — | |
| DualTalk2026.03 | 28.41 | 38.29 | 9.91 | — | 8.42 | 2.11 | 2.5 | 8.32 | 6.88 | 1.57 | 1.95 | 1.79 | — | |
| Listen-Mtype=retrieval-based, similarity=speaker motion2026.03 | 33.4 | 29.06 | 9.42 | 9.06 | 9.81 | 2.79 | 2.12 | 7.01 | 4.63 | 1.03 | 2.47 | 2.86 | 2.08 | |
| L2L2026.03 | 38.92 | 66.13 | 11.32 | — | 10.15 | 2.35 | 2.94 | 11.21 | 5.71 | 1.78 | 1.58 | 1.12 | — | |
| SelfTalk2026.03 | 43.58 | 53.98 | 8.21 | — | 11.59 | 2.47 | 2.41 | 10.98 | 6.13 | 1.68 | 1.27 | 1.39 | — | |
| CodeTalker2026.03 | 47.23 | 70.54 | 10.47 | — | 14.28 | 3.07 | 2.95 | 12.49 | 6.85 | 0 | 0 | 0 | — | |
| FaceFormer2026.03 | 52.66 | 59.84 | 13.89 | — | 12.34 | 2.96 | 2.84 | 10.47 | 6.44 | 1.59 | 0.43 | 0.86 | — | |
| DIM2026.03 | 55.09 | 45.2 | 14.56 | — | 10.67 | 2.96 | 2.35 | 11.79 | 6.21 | 0.73 | 1.84 | 1.31 | — | |
| Listen-Rtype=random listener motion baseline2026.03 | 63.74 | 68.75 | 11.03 | 8.9 | 10.93 | 2.58 | 2.27 | 7.35 | 7.98 | 1.84 | 2.39 | 2.18 | 2.98 | |
| Listen-Atype=retrieval-based, similarity=audio2026.03 | 65.09 | 41.57 | 12.57 | 7.67 | 7.98 | 2.32 | 1.91 | 7.68 | 4.89 | 1.29 | 2.26 | 2.57 | 1.96 |