Visual Reasoning on VLMs are Blind (Accuracy)
77.8AccuracyKimi-K2.5
Evaluation Results
| Method | Links | |
|---|---|---|
| Kimi-K2.5Model Category=Open-Source & Baselines2026.05 | 77.8 | |
| MiMo-VLParams=7B2025.11 | 74.91 | |
| Gemini-2.5-ProModel Category=Closed-Source Models2026.05 | 74.6 | |
| Best ModelModel Category=Open-Source & Baselines2026.05 | 73.6 | |
| MiMo-EmbodiedParams=7B2025.11 | 72.32 | |
| Claude 3.7 SonnetParams=–2025.11 | 72.1 | |
| MAESTRO*Evaluation Protocol=augments the registry with 2 additional experts and 4 new Level-1 skills2026.05 | 72.1 | |
| Qwen3-VL-32BModel Category=Open-Source & Baselines2026.05 | 71.1 | |
| GLM-4.6VModel Category=Open-Source & Baselines2026.05 | 69.6 | |
| MAESTROEvaluation Protocol=default pool2026.05 | 69.1 | |
| Gemini-2.5-FlashModel Category=Closed-Source Models2026.05 | 68.4 | |
| Direct AnsweringModel Category=Open-Source & Baselines2026.05 | 67.8 | |
| GPT-5Model Category=Closed-Source Models2026.05 | 66.7 | |
| DeepEyes-v2Model Category=Think with Images Methods2026.05 | 55.1 | |
| PixelReasonerModel Category=Think with Images Methods2026.05 | 50.9 | |
| GPT-4oParams=–2025.11 | 49.8 | |
| GPT-4oModel Category=Closed-Source Models2026.05 | 48.9 | |
| VisionReasonerModel Category=Think with Images Methods2026.05 | 48.9 | |
| Untrained ModelModel Category=Open-Source & Baselines2026.05 | 48.4 | |
| VTOOL-R1Model Category=Think with Images Methods2026.05 | 48.4 | |
| DeepEyesModel Category=Think with Images Methods2026.05 | 48.1 | |
| ThymeModel Category=Think with Images Methods2026.05 | 48.1 | |
| Visual-ARFTModel Category=Think with Images Methods2026.05 | 44.6 | |
| MathCoder-VLModel Category=Think with Images Methods2026.05 | 42.6 | |
| VTS-VModel Category=Think with Images Methods2026.05 | 42.1 | |
| Chain-of-FocusModel Category=Think with Images Methods2026.05 | 39.4 | |
| Qwen2.5-VLParams=7B2025.11 | 37.4 | |
| InternVL3Params=8B2025.11 | 36.8 |