Gaze Following on VideoAttentionTarget
0.051L2 DistanceHuman
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Human2026.05 | 0.051 | — | 0.925 | — | — | — | |
| HCLoRABackbone=ViT-L2026.05 | 0.0951 | 0.9387 | — | — | 11.75 | 90.68 | |
| OmniGFInput modality=I, Multi-person inference support=true2026.05 | 0.096 | — | 0.923 | — | — | — | |
| HCLoRABackbone=ViT-B2026.05 | 0.1006 | 0.9352 | — | — | 12.26 | 89.88 | |
| GazeVLMInput modality=I+D, Multi-person inference support=true2026.05 | 0.102 | — | 0.898 | — | — | — | |
| Gaze-LLE (ViT-L)Input modality=I, Multi-person inference support=false2026.05 | 0.103 | — | 0.903 | — | — | — | |
| Sharingan2026.05 | 0.1038 | 0.9162 | — | — | 12.09 | 89.21 | |
| Jin [27]Input modality=I+D+P, Multi-person inference support=false2026.05 | 0.104 | — | 0.895 | — | — | — | |
| GazeLLE + VPTBackbone=ViT-B, Adaptation protocol=Visual Prompt Tuning2026.05 | 0.1044 | 0.9309 | — | — | 13.66 | 89.11 | |
| MTGSInput modality=I, Multi-person inference support=true2026.05 | 0.105 | — | 0.869 | — | — | — | |
| SharinganInput modality=I, Multi-person inference support=true2026.05 | 0.107 | — | 0.891 | — | — | — | |
| Gaze-LLE (ViT-B)Input modality=I, Multi-person inference support=false2026.05 | 0.107 | — | 0.897 | — | — | — | |
| GazeLLEBackbone=ViT-B2026.05 | 0.1071 | 0.9347 | — | — | 15.02 | 89.79 | |
| GazeLLE + LoRABackbone=ViT-B, Adaptation protocol=LoRA2026.05 | 0.1077 | 0.9307 | — | — | 14.24 | 87.75 | |
| DAMHead location setting=Default2022.03 | 0.108 | 0.896 | — | — | — | — | |
| FangInput modality=I+D+E, Multi-person inference support=false2026.05 | 0.108 | — | 0.896 | — | — | — | |
| MiaoInput modality=I+D, Multi-person inference support=false2026.05 | 0.109 | — | 0.908 | — | — | — | |
| TafascaInput modality=I+D, Multi-person inference support=false2026.05 | 0.109 | — | 0.834 | — | — | — | |
| Miao et al.2026.05 | 0.1099 | 0.9163 | — | — | 12.35 | 90.61 | |
| GuptaInput modality=I+D+P, Multi-person inference support=false2026.05 | 0.11 | — | 0.879 | — | — | — | |
| GazeLLE + FTBackbone=ViT-B, Adaptation protocol=Full fine-tuning2026.05 | 0.1107 | 0.9347 | — | — | 15.32 | 88.58 | |
| BaoInput modality=I+D+P, Multi-person inference support=false2026.05 | 0.12 | — | 0.869 | — | — | — | |
| HGTTRBackbone=ResNet-101, Head location setting=Real2022.03 | 0.126 | 0.904 | 0.854 | 0.523 | — | — | |
| Jin [51]Input modality=I, Multi-person inference support=true2026.05 | 0.127 | — | 0.882 | — | — | — | |
| Chong et al.2026.05 | 0.1339 | 0.8628 | — | — | 16.14 | 85.1 | |
| VideoAttentionHead location setting=Default2022.03 | 0.134 | 0.86 | 0.853 | 0.42 | — | — | |
| Chong [12]Input modality=I, Multi-person inference support=false2026.05 | 0.134 | — | 0.853 | — | — | — | |
| HGTTRBackbone=ResNet-50, Head location setting=Real2022.03 | 0.137 | 0.893 | 0.821 | 0.514 | — | — | |
| VideoAttentionHead location setting=Real2022.03 | 0.146 | 0.812 | 0.849 | 0.42 | — | — | |
| LianHead location setting=Default2022.03 | 0.165 | 0.837 | — | 0.392 | — | — | |
| Chong [22]Input modality=I, Multi-person inference support=false2026.05 | 0.171 | — | 0.712 | — | — | — | |
| LianHead location setting=Real2022.03 | 0.172 | 0.784 | — | 0.392 | — | — | |
| ChongHead location setting=Default2022.03 | 0.193 | 0.83 | 0.705 | 0.374 | — | — | |
| ChongHead location setting=Real2022.03 | 0.214 | 0.791 | 0.651 | 0.374 | — | — | |
| Fixed biasHead location setting=Default2022.03 | 0.326 | 0.728 | 0.624 | 0.13 | — | — | |
| RandomHead location setting=Default2022.03 | 0.458 | 0.505 | 0.621 | 0.091 | — | — | |
| Fixed biasHead location setting=Real2022.03 | 0.472 | 0.522 | 0.51 | 0.13 | — | — | |
| RandomHead location setting=Real2022.03 | 0.592 | 0.247 | 0.349 | 0.091 | — | — |