Multimodal Reasoning on ERQA
55AccuracyBest Model
Evaluation Results
| Method | Links | |
|---|---|---|
| Best ModelModel Category=Open-Source & Baselines2026.05 | 55 | |
| MAESTRO*Evaluation Protocol=augments the registry with 2 additional experts and 4 new Level-1 skills2026.05 | 52.5 | |
| Qwen3-VL-32BModel Category=Open-Source & Baselines2026.05 | 49.3 | |
| Direct AnsweringModel Category=Open-Source & Baselines2026.05 | 47.5 | |
| ThymeModel Category=Think with Images Methods2026.05 | 47.3 | |
| Kimi-K2.5Model Category=Open-Source & Baselines2026.05 | 46 | |
| GPT-5Model Category=Closed-Source Models2026.05 | 45.8 | |
| GPT-4oModel Category=Closed-Source Models2026.05 | 44.8 | |
| Gemini-2.5-FlashModel Category=Closed-Source Models2026.05 | 44.3 | |
| GLM-4.6VModel Category=Open-Source & Baselines2026.05 | 44 | |
| Gemini-2.5-ProModel Category=Closed-Source Models2026.05 | 43 | |
| MAESTROEvaluation Protocol=default pool2026.05 | 42.8 | |
| DeepEyes-v2Model Category=Think with Images Methods2026.05 | 42.3 | |
| Chain-of-FocusModel Category=Think with Images Methods2026.05 | 41.8 | |
| Untrained ModelModel Category=Open-Source & Baselines2026.05 | 41.5 | |
| VisionReasonerModel Category=Think with Images Methods2026.05 | 39.3 | |
| VTS-VModel Category=Think with Images Methods2026.05 | 37.5 | |
| MathCoder-VLModel Category=Think with Images Methods2026.05 | 37.3 | |
| Visual-ARFTModel Category=Think with Images Methods2026.05 | 37 | |
| VTOOL-R1Model Category=Think with Images Methods2026.05 | 36.8 | |
| DeepEyesModel Category=Think with Images Methods2026.05 | 35.5 | |
| PixelReasonerModel Category=Think with Images Methods2026.05 | 33.5 |