Visual Spatial Intelligence Reasoning on VSI-Debiased
0.699AccuracySSR-3D
Evaluation Results
| Method | Links | |
|---|---|---|
| SSR-3DCategory=Open-Sourced spatial models2026.02 | 0.699 | |
| GeoThinker Qwen3vl-8B-32frame# Frames=1282026.02 | 0.681 | |
| GeoThinker Qwen3vl-8B-32frame# Frames=642026.02 | 0.677 | |
| SSR-2DCategory=Open-Sourced spatial models2026.02 | 0.666 | |
| GeoThinker Qwen3vl-8B-32frame# Frames=322026.02 | 0.663 | |
| GeoThinker Qwen3vl-8B-8frame# Frames=1282026.02 | 0.653 | |
| GeoThinker Qwen3vl-8B-8frame# Frames=322026.02 | 0.648 | |
| GeoThinker Qwen3vl-8B-8frame# Frames=642026.02 | 0.643 | |
| GeoThinker Qwen3vl-8B-32frame# Frames=162026.02 | 0.643 | |
| SenseNova-SI InternVL3-8B# Frames=322026.02 | 0.628 | |
| SenseNova-SI (InternVL3-8B)Category=Open-Sourced spatial models2026.02 | 0.628 | |
| SenseNova-SI InternVL3-8B# Frames=642026.02 | 0.624 | |
| GeoThinker Qwen3vl-8B-8frame# Frames=162026.02 | 0.607 | |
| Cambrian-S-7B# Frames=1282026.02 | 0.599 | |
| SenseNova-SI InternVL3-8B# Frames=1282026.02 | 0.597 | |
| Cambrian-S-7B# Frames=642026.02 | 0.591 | |
| SenseNova-SI InternVL3-8B# Frames=162026.02 | 0.589 | |
| Cambrian-S-7BCategory=Open-Sourced spatial models2026.02 | 0.563 | |
| Cambrian-S-7B# Frames=322026.02 | 0.556 | |
| VG-LLM-8B*# Frames=642026.02 | 0.552 | |
| VG-LLM-8B*# Frames=1282026.02 | 0.551 | |
| VG-LLM-8B*# Frames=322026.02 | 0.524 | |
| VG-LLM-8B*# Frames=162026.02 | 0.516 | |
| Cambrian-S-7B# Frames=162026.02 | 0.497 | |
| LLaVA-Video-7BCategory=Open-Sourced general models2026.02 | 0.307 | |
| LLaVA-OneVision-7BCategory=Open-Sourced general models2026.02 | 0.285 | |
| SmolVLM2-2.2BCategory=Open-Sourced general models2026.02 | 0.223 |