Math Reasoning on MathVista (Score)
78.8ScoreVisPlay-8B-iter3
Evaluation Results
| Method | Links | |
|---|---|---|
| VisPlay-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 78.8 | |
| Jigsaw-R1-8BModel Category=Template-based Self-evolution Methods, Scale=8B2026.04 | 77.9 | |
| EVE (Ours-8B-iter4)Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=42026.04 | 77.9 | |
| Qwen3-VL-8B-InstructModel Category=Open-Source MLLMs, Scale=8B2026.04 | 76.9 | |
| HEEDMethod ID=C4, Distillation Strategy=HEED, Training Stage=Before SFT+DPO2026.05 | 76.2 | |
| TeacherMethod ID=C0, Distillation Strategy=Teacher, Training Stage=Before SFT+DPO2026.05 | 76 | |
| RSAMethod ID=C3, Distillation Strategy=RSA, Training Stage=Before SFT+DPO2026.05 | 75.5 | |
| HSAMethod ID=C2, Distillation Strategy=HSA, Training Stage=Before SFT+DPO2026.05 | 75.4 | |
| ThinkLite-VLModel Category=Multimodal Reasoning Models2026.03 | 75.1 | |
| KDMethod ID=C1, Distillation Strategy=KD, Training Stage=Before SFT+DPO2026.05 | 74.9 | |
| AVAR-ThinkerModel Category=Our model2026.03 | 74.7 | |
| MM-Zero-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 74.7 | |
| Claude-3.7-SonnetModel Category=Closed-Source2026.03 | 74.5 | |
| Vision-R1Model Category=Multimodal Reasoning Models, Trained on MathVision=True2026.03 | 73.5 | |
| MM-Eureka-7BModel Category=Multimodal Reasoning Models2026.03 | 73 | |
| OpenVLThinkerModel Category=Multimodal Reasoning Models2026.03 | 72.3 | |
| GPT-5 nano (high)Model Category=Closed-Source MLLMs, Scale=high2026.04 | 71.5 | |
| InternVL3-9BModel Category=Open-Source MLLMs, Scale=9B2026.04 | 71.5 | |
| Qwen2.5-VL-7BModel Category=Open-Source General Models2026.03 | 68.2 | |
| Vision-SR1Model Category=Multimodal Reasoning Models2026.03 | 68.1 | |
| VLAA-Thinker-7BModel Category=Multimodal Reasoning Models2026.03 | 68 | |
| LLaVA-OneVision-72BModel Category=Open-Source MLLMs, Scale=72B2026.04 | 67.1 | |
| InternVL2.5-8BModel Category=Open-Source General Models2026.03 | 64.4 | |
| R1-OneVisionModel Category=Multimodal Reasoning Models2026.03 | 64.1 | |
| GPT-4oModel Category=Closed-Source2026.03 | 63.8 | |
| GPT-4o-20240513Model Category=Closed-Source MLLMs2026.04 | 63.8 | |
| Mulberry-7BModel Category=Multimodal Reasoning Models, Trained on MathVision=True2026.03 | 63.1 | |
| GPT-5 mini (minimal)Model Category=Closed-Source MLLMs, Scale=minimal2026.04 | 59.6 | |
| LLaVA-OneVision-7BModel Category=Open-Source General Models2026.03 | 58.6 | |
| Llama-3.2-11B-Vision-InstructModel Category=Open-Source General Models2026.03 | 48.6 |