Multiple Choice Answering on VIEW2SPACE
93.57Accuracy (%)Human-max
Evaluation Results
| Method | Links | |
|---|---|---|
| Human-max2026.03 | 93.57 | |
| Human-avg2026.03 | 89.88 | |
| Human-min2026.03 | 85 | |
| Ours (Grounded CoT)Model Type=Fine-tuned Model, Backbone=Qwen3VL-4B2026.03 | 64.93 | |
| GPT-52026.03 | 59.86 | |
| Qwen3-VL-4B2026.03 | 35.19 | |
| Random (chance)2026.03 | 28.59 | |
| Random (frequency)2026.03 | 28.59 |