Mathematical Visual Question Answering on MathVerse
82.9AccuracyGemini 2.5 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 2.5 Pro2026.04 | 82.9 | |
| OpenVLThinkerV2Parameters=8B2026.04 | 65.8 | |
| Qwen3-VL GDPORL Strategy=GDPO2026.04 | 64.9 | |
| Qwen3-VL GRPORL Strategy=GRPO2026.04 | 64.7 | |
| OneThinker-8BParameters=8B2026.04 | 64.3 | |
| InternVL3.5-4BFrames=-2026.03 | 61.7 | |
| InternVL3.5-8BFrames=-2026.03 | 61.5 | |
| Qwen3-VL-8B-InstructFrames=642026.03 | 58.1 | |
| Qwen3-VL-Instruct-8BParameters=8B2026.04 | 58.1 | |
| ARES-7BParameters=7B2026.04 | 56.5 | |
| R-MSD (2B)Frames=642026.03 | 55.3 | |
| OVR-7BParameters=7B2026.04 | 54.6 | |
| VL-Rethinker-7BParameters=7B2026.04 | 54.2 | |
| InternVL3.5-2BFrames=-2026.03 | 53.4 | |
| VisionZero2026.04 | 52.1 | |
| Qwen3-VL-2B-InstructFrames=642026.03 | 51.1 | |
| MM-Eureka-7BParameters=7B2026.04 | 50.3 | |
| Vision-G12026.04 | 50 | |
| R-MSD (4B)Frames=642026.03 | 49.3 | |
| OpenVLThinker-7BParameters=7B2026.04 | 47.9 | |
| Original SFT+RL (4B)Frames=642026.03 | 46.8 | |
| Qwen3-VL-4B-InstructFrames=642026.03 | 45.7 | |
| GPT-4oFrames=-2026.03 | 41.2 | |
| GPT-4o2026.04 | 41.2 | |
| InternVL3-8BFrames=-2026.03 | 39.8 | |
| InternVL2.5-8BFrames=-2026.03 | 39.5 | |
| InternVL2.5-4BFrames=-2026.03 | 37.1 | |
| InternVL2.5-2BFrames=-2026.03 | 30.6 | |
| MaD-MixModel size=7B, Evaluation protocol=0-shot2026.02 | 26.4 | |
| UniformModel size=7B, Evaluation protocol=0-shot2026.02 | 25.63 | |
| AvgModel size=7B, Evaluation protocol=0-shot2026.02 | 25.52 | |
| InternVL3-2BFrames=-2026.03 | 25.3 | |
| FusedModel size=7B, Evaluation protocol=0-shot2026.02 | 24.75 | |
| FusedModel size=0.5B, Evaluation protocol=0-shot2026.02 | 17.77 | |
| AvgModel size=0.5B, Evaluation protocol=0-shot2026.02 | 15.62 | |
| UniformModel size=0.5B, Evaluation protocol=0-shot2026.02 | 15.61 | |
| MaD-MixModel size=0.5B, Evaluation protocol=0-shot2026.02 | 15.1 |