Egocentric visual attention prediction on Ego4D (test)
0.401F1 ScoreLanguage-guided scene context-aware learning framework
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Language-guided scene context-aware learning frameworkAdditional sensory modality=video frame only2026.01 | 0.401 | 0.541 | 0.319 | |
| CSTSAdditional sensory modality=audio2026.01 | 0.397 | 0.533 | 0.316 | |
| GLCAdditional sensory modality=video frame only2026.01 | 0.378 | 0.529 | 0.294 | |
| DFG+Additional sensory modality=video frame only2026.01 | 0.373 | 0.523 | 0.29 | |
| DFGAdditional sensory modality=video frame only2026.01 | 0.372 | 0.532 | 0.286 | |
| MViTAdditional sensory modality=video frame only2026.01 | 0.372 | 0.541 | 0.283 | |
| AttnTransAdditional sensory modality=flow2026.01 | 0.37 | 0.55 | 0.279 | |
| I3D-R50Additional sensory modality=video frame only2026.01 | 0.369 | 0.521 | 0.286 | |
| GazeMLEAdditional sensory modality=flow2026.01 | 0.363 | 0.525 | 0.278 |