Urban Perception Prediction on Urban Perception Boring (test)
45.5Macro-F1Gaze + Patch Transformer
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gaze + Patch TransformerRepresentation=Pretrained ViT patch, Gaze Integration=Explicit gaze token features2026.05 | 45.5 | 46.6 | |
| Gaze + Patch Transformer (Shuffled Patch Alignment)Ablation=Random alignment between gaze and patches2026.05 | 44.9 | 46.4 | |
| Image-Only ViTRepresentation=Pretrained ViT patch2026.05 | 44.8 | 46.2 | |
| Gaze + Patch Transformer (w/o Gaze)Ablation=No gaze input2026.05 | 44.5 | 46.3 | |
| Gaze-weighted Patch PoolingRepresentation=Pretrained ViT patch, Fusion=Soft spatial prior reweighting2026.05 | 44 | 45 | |
| Patch Sequence TransformerRepresentation=Pretrained ViT patch, Gaze Integration=Sequence order only2026.05 | 43.3 | 44.9 |