Multimodal Mathematical Reasoning on OlympiadBench
68Accuracyo1
Evaluation Results
| Method | Links | |
|---|---|---|
| o12026.03 | 68 | |
| o1Model Category=Closed-Source Models2026.03 | 68 | |
| Gemini2-flash2026.03 | 51 | |
| Gemini-2-flashModel Category=Closed-Source Models2026.03 | 51 | |
| Claude3.7-Sonnet2026.03 | 48.9 | |
| Claude3.7-SonnetModel Category=Closed-Source Models2026.03 | 48.9 | |
| MM-Eureka-32B-R-TAPModel Category=Open-Source Reasoning Models2026.03 | 41.2 | |
| Qwen-2.5-VL-72B2026.03 | 40.4 | |
| Qwen-2.5-VL-72BModel Category=Open-Source General Models2026.03 | 40.4 | |
| MM-Eureka-32BModel Category=Open-Source Reasoning Models2026.03 | 35.9 | |
| Claude3.7-SonnetParam (B)=-2025.09 | 35.2 | |
| GPT-4o2026.03 | 35 | |
| GPT-4oModel Category=Closed-Source Models2026.03 | 35 | |
| QVQ-72B-Preview2026.03 | 33.2 | |
| QVQ-72B-PreviewModel Category=Open-Source Reasoning Models2026.03 | 33.2 | |
| InternVL2.5-VL-38B2026.03 | 32 | |
| InternVL2.5-VL-38BModel Category=Open-Source General Models2026.03 | 32 | |
| InternVL2.5-VL-78B2026.03 | 31.1 | |
| InternVL2.5-VL-78BModel Category=Open-Source General Models2026.03 | 31.1 | |
| Qwen-2.5-VL-32B2026.03 | 30 | |
| Qwen-2.5-VL-32BModel Category=Open-Source General Models2026.03 | 30 | |
| GPT-4.1-20250414Param (B)=-2025.09 | 29.4 | |
| MM-Eureka-7B-R-TAPModel Category=Open-Source Reasoning Models2026.03 | 27.5 | |
| GPT-4oParam (B)=-2025.09 | 25.9 | |
| InternVL2.5-38B-MPO2026.03 | 25.6 | |
| InternVL2.5-38B-MPOModel Category=Open-Source Reasoning Models2026.03 | 25.6 | |
| DIVA-GRPO-7B2026.03 | 23.1 | |
| Qwen2.5VL-72B x Qwen3-32BParam (B)=1042025.09 | 22.4 | |
| R1-ShareVL-7B2026.03 | 21.3 | |
| Qwen-2.5-VL-7B2026.03 | 20.2 | |
| Qwen-2.5-VL-7BModel Category=Open-Source General Models2026.03 | 20.2 | |
| QVQ-72B-PreviewParam (B)=722025.09 | 20.2 | |
| Adora-7B2026.03 | 20.1 | |
| OpenVLThinker-7B2026.03 | 20.1 | |
| MM-Eureka-7B2026.03 | 20.1 | |
| ADORA-7BModel Category=Open-Source Reasoning Models2026.03 | 20.1 | |
| OpenVLThinker-7BModel Category=Open-Source Reasoning Models2026.03 | 20.1 | |
| MM-Eureka-7BModel Category=Open-Source Reasoning Models2026.03 | 20.1 | |
| OpenVLThinker-7BParam (B)=72025.09 | 20.1 | |
| Qwen2.5-VL-72BParam (B)=722025.09 | 19.9 | |
| R1-OnevisionParam (B)=72025.09 | 17.8 | |
| Qwen2.5VL-7B x Qwen3-32BParam (B)=392025.09 | 17.4 | |
| R1-Onevision-7B2026.03 | 17.3 | |
| R1-Onevision-7BModel Category=Open-Source Reasoning Models2026.03 | 17.3 | |
| SFT-7B2026.03 | 16.2 | |
| InternVL3.5-8BParam (B)=82025.09 | 16.1 | |
| Keye-VL-8BParam (B)=82025.09 | 15.4 | |
| Qwen2.5VL-7B x Qwen3-4BParam (B)=112025.09 | 14.7 | |
| SAIL-VL2-8BParam (B)=82025.09 | 14.1 | |
| InternVL2.5-VL-8B2026.03 | 12.3 | |
| InternVL2.5-VL-8BModel Category=Open-Source General Models2026.03 | 12.3 | |
| Qwen2.5-VL-7BParam (B)=72025.09 | 8.6 | |
| InternVL2.5-8B-MPO2026.03 | 7.8 | |
| InternVL2.5-8B-MPOModel Category=Open-Source Reasoning Models2026.03 | 7.8 | |
| VLM-R1-3B-MathParam (B)=32025.09 | 7.3 | |
| Qwen2.5-VL-3BParam (B)=32025.09 | 7 |