Vision-Language Reasoning on SQA3D ScanNet scenes (test)
54.9BLEU-1DINOv3 + SpatialBoost
Evaluation Results
| Method | Links | |
|---|---|---|
| DINOv3 + SpatialBoostBase Model=DINOv3, SpatialBoost=true2026.03 | 54.9 | |
| PE-CoreEncoder Type=Vision-Language trained encoder2026.03 | 51.7 | |
| DINOv32026.03 | 51.4 | |
| dino.txtEncoder Type=Vision-Language trained encoder2026.03 | 50.4 | |
| DINOv2 + SpatialBoostBase Model=DINOv2, SpatialBoost=true2026.03 | 50.4 | |
| SigLIPv2 + SpatialBoostBase Model=SigLIPv2, SpatialBoost=true2026.03 | 50.1 | |
| OpenCLIP + SpatialBoostBase Model=OpenCLIP, SpatialBoost=true2026.03 | 49.9 | |
| DINOv22026.03 | 49.8 | |
| TIPSEncoder Type=Vision-Language trained encoder2026.03 | 49.2 | |
| V-JEPAv2Encoder Type=Vision-only trained encoder2026.03 | 49 | |
| SigLIPv22026.03 | 48.5 | |
| AIMv2Encoder Type=Vision-Language trained encoder2026.03 | 48.1 | |
| OpenCLIP2026.03 | 48 |