Visual Attention Prediction on Aria Everyday Activities (AEA) unseen (test)
53.7F1 Scorelanguage-guided scene context-aware learning framework
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| language-guided scene context-aware learning frameworkZero-shot=true, Modality=Video frame only2026.01 | 53.7 | 65 | 45.7 | |
| CSTSZero-shot=true, Modality=Audio2026.01 | 50.8 | 62.2 | 42.9 | |
| GLCZero-shot=true, Modality=Video frame only2026.01 | 46.9 | 72.8 | 34.6 | |
| MViTZero-shot=true, Modality=Video frame only2026.01 | 44.1 | 59.7 | 35 | |
| GazeMLEZero-shot=true, Modality=Optical flow2026.01 | 44 | 59 | 35 | |
| AttnTransZero-shot=true, Modality=Optical flow2026.01 | 43.1 | 57.5 | 34.5 | |
| DFG+Zero-shot=true, Modality=Video frame only2026.01 | 43.1 | 76.4 | 30 | |
| I3D-R50Zero-shot=true, Modality=Video frame only2026.01 | 41.5 | 77.2 | 28.4 | |
| DFGZero-shot=true, Modality=Video frame only2026.01 | 39.3 | 80.4 | 26 |