Multimodal Reasoning on Vstar Bench Spatial
90.8AccuracyDeepconf
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepconfModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 90.8 | 55 | |
| RTWIModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 90.8 | 58.5 | |
| SCModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 89.5 | — | |
| ASCModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 89.5 | 56.8 | |
| ESCModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 89.5 | 50.9 | |
| CISCModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 89.5 | 56.6 | |
| Self-Cer.Model Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 89.5 | 57 | |
| DeepEyesOnline setting=true2026.02 | 88.2 | — | |
| BaseModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 86.8 | — | |
| RTWIModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 81.6 | 61.2 | |
| ThymeOnline setting=true2026.02 | 80.3 | — | |
| ASCModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 79 | 42.8 | |
| SCModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 77.6 | — | |
| ESCModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 77.6 | 22 | |
| DeepconfModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 76.3 | 52.8 | |
| BaseModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 73.7 | — | |
| CISCModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 69.7 | 29.9 | |
| Self-Cer.Model Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 68.4 | 33.1 | |
| GPT-4oOnline setting=true2026.02 | 60.5 | — |