Urban Perception Prediction on Urban Perception Safe (test)
50.5Macro F1Gaze + Patch Transformer
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gaze + Patch TransformerRepresentation=Pretrained ViT patch, Gaze Integration=Explicit gaze token features2026.05 | 50.5 | 52.1 | |
| Gaze + Patch Transformer (w/o Gaze)Ablation=No gaze input2026.05 | 50.1 | 51.7 | |
| Patch Sequence TransformerRepresentation=Pretrained ViT patch, Gaze Integration=Sequence order only2026.05 | 49.9 | 51.9 | |
| Gaze + Patch Transformer (Shuffled Patch Alignment)Ablation=Random alignment between gaze and patches2026.05 | 49.9 | 50.9 | |
| Image-Only ViTRepresentation=Pretrained ViT patch2026.05 | 49.1 | 50.7 | |
| Gaze-weighted Patch PoolingRepresentation=Pretrained ViT patch, Fusion=Soft spatial prior reweighting2026.05 | 48.8 | 50.9 | |
| Gaze + AOI TransformerInput=Gaze + Semantic AOI2026.05 | 42.2 | 45 | |
| AOI Sequence TransformerInput=Semantic AOI sequence2026.05 | 39.6 | 43.2 | |
| w/o GazeAblation=Removal of gaze component from Gaze + AOI Transformer2026.05 | 39.4 | 43.4 | |
| AOI-Only ModelInput=Semantic AOI only2026.05 | 39.2 | 42.1 | |
| Shuffled AOI AlignmentAblation=Randomized alignment for Gaze + AOI Transformer2026.05 | 36.6 | 39.5 |