Multimodal Reasoning on MathVerse MINI
77.7AccuracyQwen3-VL-8B-Thinking
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL-8B-ThinkingModel Type=Reasoning2025.12 | 77.7 | |
| Qwen3-VL-4B-ThinkingModel Type=Reasoning2025.12 | 75.2 | |
| MiMo-VL-7B-RLModel Type=Reasoning2025.12 | 71.5 | |
| Qwen3-VL-8B-Instruct + CAREBackbone=Qwen3-VL-8B-Instruct, Method=CARE2025.12 | 69.7 | |
| Gemini-2.0-ProModel Type=Proprietary2025.12 | 67.3 | |
| MiMo-VL-7B-SFTModel Type=Instruct2025.12 | 67.1 | |
| SwimBird2026.02 | 65.8 | |
| Qwen3-VL-8B-InstructModel Type=Instruct2025.12 | 62.1 | |
| Qwen3-VL-8B-Instructreproduced by authors=true2026.02 | 61.3 | |
| Keye-VL-1.5-8BModel Type=Reasoning2025.12 | 59.8 | |
| Qwen2.5-VL-7B + CAREBackbone=Qwen2.5-VL-7B, Method=CARE2025.12 | 56.8 | |
| Qwen2.5-VL-7B + GSPOBackbone=Qwen2.5-VL-7B, RL Method=GSPO2025.12 | 56 | |
| Qwen2.5-VL-7B + DAPOBackbone=Qwen2.5-VL-7B, RL Method=DAPO2025.12 | 54.2 | |
| DeepEyesV22026.02 | 52.7 | |
| Qwen2.5-VL-7B + GRPOBackbone=Qwen2.5-VL-7B, RL Method=GRPO2025.12 | 50.8 | |
| Qwen2.5-VL-3B + CAREBackbone=Qwen2.5-VL-3B, Method=CARE2025.12 | 49.9 | |
| Qwen2.5-VL-7BModel Type=Instruct2025.12 | 49.2 | |
| Qwen2.5-VL-32B-Instruct2026.02 | 48.5 | |
| Qwen2.5-VL-3BModel Type=Instruct2025.12 | 47.6 | |
| DeepEyes2026.02 | 47.3 | |
| DeepEyesModel Type=Reasoning2025.12 | 47.3 | |
| Qwen3-VL-4B-InstructModel Type=Instruct2025.12 | 46.8 | |
| Qwen2.5-VL-7B-Instruct2026.02 | 45.6 | |
| GPT-4oModel Type=Proprietary2025.12 | 37.6 | |
| LLaVA-OneVision2026.02 | 19.3 |