Gaze Following on GazeFollowing
0.04Minimum DistanceOmniGF
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OmniGFInput modality=I, Multi-person inference support=true2026.05 | 0.04 | — | 0.091 | — | |
| Human2026.05 | 0.04 | — | 0.096 | — | |
| Gaze-LLE (ViT-L)Input modality=I, Multi-person inference support=false2026.05 | 0.041 | — | 0.099 | — | |
| Gaze-LLE (ViT-B)Input modality=I, Multi-person inference support=false2026.05 | 0.045 | — | 0.104 | — | |
| GuptaInput modality=I+D+P, Multi-person inference support=false2026.05 | 0.056 | — | 0.114 | — | |
| SharinganInput modality=I, Multi-person inference support=true2026.05 | 0.057 | — | 0.113 | — | |
| MTGSInput modality=I, Multi-person inference support=true2026.05 | 0.059 | — | 0.116 | — | |
| GazeVLMInput modality=I+D, Multi-person inference support=true2026.05 | 0.061 | — | 0.123 | — | |
| TafascaInput modality=I+D, Multi-person inference support=false2026.05 | 0.062 | — | 0.122 | — | |
| Jin [27]Input modality=I+D+P, Multi-person inference support=false2026.05 | 0.063 | — | 0.118 | — | |
| HGTTRBackbone=ResNet-101, Head location setting=Real2022.03 | 0.065 | 0.905 | 0.138 | 54.1 | |
| MiaoInput modality=I+D, Multi-person inference support=false2026.05 | 0.065 | — | 0.123 | — | |
| DAMHead location setting=Default2022.03 | 0.067 | 0.922 | 0.124 | — | |
| FangInput modality=I+D+E, Multi-person inference support=false2026.05 | 0.067 | — | 0.124 | — | |
| HGTTRBackbone=ResNet-50, Head location setting=Real2022.03 | 0.069 | 0.917 | 0.133 | 54.7 | |
| Jin [51]Input modality=I, Multi-person inference support=true2026.05 | 0.076 | — | 0.126 | — | |
| VideoAttentionHead location setting=Default2022.03 | 0.077 | 0.921 | 0.137 | 48.3 | |
| Chong [12]Input modality=I, Multi-person inference support=false2026.05 | 0.077 | — | 0.137 | — | |
| LianHead location setting=Default2022.03 | 0.081 | 0.906 | 0.145 | 46.9 | |
| LianInput modality=I, Multi-person inference support=false2026.05 | 0.081 | — | 0.145 | — | |
| ZhaoHead location setting=Real2022.03 | 0.082 | — | — | — | |
| VideoAttentionHead location setting=Real2022.03 | 0.082 | 0.902 | 0.142 | 48.3 | |
| LianHead location setting=Real2022.03 | 0.087 | 0.881 | 0.153 | 46.9 | |
| ChongHead location setting=Default2022.03 | 0.112 | 0.896 | 0.187 | 44.9 | |
| Chong [22]Input modality=I, Multi-person inference support=false2026.05 | 0.112 | — | 0.187 | — | |
| GazeFollowHead location setting=Default2022.03 | 0.113 | 0.878 | 0.19 | 45.7 | |
| RecasensInput modality=I, Multi-person inference support=false2026.05 | 0.113 | — | 0.19 | — | |
| ChongHead location setting=Real2022.03 | 0.12 | 0.807 | 0.207 | 44.9 | |
| GazeFollowHead location setting=Real2022.03 | 0.124 | 0.804 | 0.233 | 45.7 | |
| ZhaoHead location setting=Default2022.03 | 0.147 | — | — | — | |
| CenterHead location setting=Default2022.03 | 0.23 | 0.633 | 0.313 | 11.7 | |
| JuddHead location setting=Default2022.03 | 0.25 | 0.711 | 0.337 | — | |
| CenterHead location setting=Real2022.03 | 0.371 | 0.446 | 0.495 | 11.7 | |
| RandomHead location setting=Default2022.03 | 0.391 | 0.504 | 0.484 | 10.4 | |
| RandomHead location setting=Real2022.03 | 0.487 | 0.381 | 0.533 | 10.4 | |
| BaoInput modality=I+D+P, Multi-person inference support=false2026.05 | — | — | 0.122 | — |