Mathematical Reasoning on DynaMath
81.42AccuracyGemini2.5-Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini2.5-ProLLM Solver=Qwen3-30A3-Thinking2026.03 | 81.42 | |
| GPT-5 miniVersion=high2026.05 | 81.4 | |
| Official thinkingSource=Qwen3-VL2026.05 | 80.1 | |
| RLVRTraining Source=DeepVision2026.05 | 80 | |
| CodePercept-32B-S1LLM Solver=Qwen3-235A22-Thinking2026.03 | 79.5 | |
| RLVRTraining Source=OpenMMR2026.05 | 78.7 | |
| RLR³Training Source=ViRL2026.05 | 78.4 | |
| RLR³Training Source=OpenMMR2026.05 | 78 | |
| CodePercept-32B-S1LLM Solver=Qwen3-30A3-Thinking2026.03 | 77.41 | |
| Qwen3-VL-235A22B-InstructLLM Solver=Qwen3-30A3-Thinking2026.03 | 77.39 | |
| RLR³Training Source=DeepVision2026.05 | 76.4 | |
| Qwen3-VL-32B-InstructLLM Solver=Qwen3-30A3-Thinking2026.03 | 75.78 | |
| Qwen3-VL-32B-InstructLLM Solver=Qwen3-235A22-Thinking2026.03 | 75.54 | |
| CodePercept-8B-S1LLM Solver=Qwen3-235A22-Thinking2026.03 | 75.05 | |
| Base instruct2026.05 | 74.8 | |
| RLVRTraining Source=ViRL2026.05 | 74.1 | |
| Qwen3-VL-8B-InstructLLM Solver=Qwen3-235A22-Thinking2026.03 | 73.69 | |
| Official instructSource=Qwen3-VL2026.05 | 73.4 | |
| Claude-Opus 4.1-ThinkingLLM Solver=Qwen3-30A3-Thinking2026.03 | 73.25 | |
| CodePercept-8B-S1LLM Solver=Qwen3-30A3-Thinking2026.03 | 73.2 | |
| CodePercept-4B-S1LLM Solver=Qwen3-235A22-Thinking2026.03 | 72.4 | |
| Qwen3-VL-8B-InstructLLM Solver=Qwen3-30A3-Thinking2026.03 | 72.19 | |
| Qwen3-VL-30A3B-InstructLLM Solver=Qwen3-30A3-Thinking2026.03 | 71.67 | |
| CodePercept-4B-S1LLM Solver=Qwen3-30A3-Thinking2026.03 | 71.38 | |
| GPT-5 miniVersion=minimal2026.05 | 71.3 | |
| Qwen3-VL-4B-InstructLLM Solver=Qwen3-235A22-Thinking2026.03 | 71.22 | |
| GPT5-ThinkingLLM Solver=Qwen3-30A3-Thinking2026.03 | 71 | |
| OPD+ViCuRModel Scale=8B2026.06 | 70.7 | |
| Qwen3-VL-4B-Instruct / H-OPD (8B)Teacher=8B Ensemble, Method=H-OPD2026.07 | 69.9 | |
| Qwen3-VL-4B-InstructLLM Solver=Qwen3-30A3-Thinking2026.03 | 69.4 | |
| OPDModel Scale=8B2026.06 | 69.22 | |
| Qwen2.5-VL-72BLLM Solver=Qwen3-30A3-Thinking2026.03 | 68.28 | |
| InternVL3.5-8BLLM Solver=Qwen3-30A3-Thinking2026.03 | 68.12 | |
| Qwen3-VL-8B2026.07 | 67.7 | |
| GRPOModel Scale=8B2026.06 | 67.44 | |
| OPSD+ViCuRModel Scale=8B2026.06 | 67.17 | |
| Base ModelModel Scale=8B2026.06 | 67.13 | |
| GLM-4.1V-9BLLM Solver=Qwen3-30A3-Thinking2026.03 | 66.17 | |
| OPSDModel Scale=8B2026.06 | 65.99 | |
| MiniCPM-V-4.5LLM Solver=Qwen3-30A3-Thinking2026.03 | 65.44 | |
| Qwen3-VL-4B2026.07 | 65.3 | |
| Claude-3.5 SonnetActivation Replay=false2025.11 | 64.8 | |
| GPT-4oActivation Replay=false2025.11 | 63.7 | |
| Intern-S1-8BLLM Solver=Qwen3-30A3-Thinking2026.03 | 63.61 | |
| KeyeVL1.5-8BLLM Solver=Qwen3-30A3-Thinking2026.03 | 62.37 | |
| MM-Eureka-Qwen-32BActivation Replay=false2025.11 | 62.1 | |
| MM-Eureka-Qwen-32BActivation Replay=true2025.11 | 61.8 | |
| Qwen3-VL-2B-Instruct / H-OPD (8B)Teacher=8B Ensemble, Method=H-OPD2026.07 | 59.5 | |
| AutoToolSize=7B2026.05 | 58 | |
| Qwen3-VL-2B-Instruct / H-OPD (4B)Teacher=4B Ensemble, Method=H-OPD2026.07 | 57.9 | |
| InternVL3-8BSize=8B2026.05 | 57.8 | |
| Qwen2.5-VL-7B + PGPOModel Size=7B, Optimization Strategy=PGPO2026.04 | 57.71 | |
| DeepEyesSize=7B2026.05 | 57.7 | |
| Qwen2.5-VL-7BSize=7B2026.05 | 57.2 | |
| Qwen2.5-VL-7B + VPPOModel Size=7B, Optimization Strategy=VPPO2026.04 | 57.11 | |
| Qwen2.5-VL-7B + PAPOModel Size=7B, Optimization Strategy=PAPO2026.04 | 56.19 | |
| Qwen2.5-VL-7B + DAPOModel Size=7B, Optimization Strategy=DAPO2026.04 | 55.96 | |
| Qwen3-VL-2B-Instruct / GRPOMethod=GRPO2026.07 | 55.7 | |
| GRPOModel Scale=2B2026.06 | 55.49 | |
| MM-Eureka-7BModel Size=7B2026.04 | 55.01 | |
| DeepEyesParam Size=7B2025.05 | 55 | |
| VL-Rethinker-7BModel Size=7B2026.04 | 54.97 | |
| VL-Rethinker-7BActivation Replay=true2025.11 | 54.9 | |
| NoisyRollout-7BModel Size=7B2026.04 | 54.89 | |
| Qwen2.5-VL-7B + GRPOModel Size=7B, Optimization Strategy=GRPO2026.04 | 54.84 | |
| R1-ShareVL-7BModel Size=7B2026.04 | 54.8 | |
| VL-Rethinker-7BActivation Replay=false2025.11 | 54.7 | |
| MMR1-Math-v0-7BActivation Replay=true2025.11 | 54.2 | |
| Qwen3-VL-2B2026.07 | 54.2 | |
| MMR1-Math-v0-7BActivation Replay=false2025.11 | 53.8 | |
| Qwen2.5-VL*Param Size=7B2025.05 | 53.3 | |
| PLM-HoneyBee-8BModel scale=7B-8B2025.10 | 53.3 | |
| Base ModelModel Scale=2B2026.06 | 53.23 | |
| Qwen2.5-VL-7BActivation Replay=false2025.11 | 53.2 | |
| OPDModel Scale=2B2026.06 | 52.87 | |
| MM-Eureka-Qwen-7BActivation Replay=true2025.11 | 52.4 | |
| OPD+ViCuRModel Scale=2B2026.06 | 52.38 | |
| OPSDModel Scale=2B2026.06 | 52.37 | |
| PLM-HoneyBee-3BModel scale=3B-4B2025.10 | 51.9 | |
| MM-Eureka-Qwen-7BActivation Replay=false2025.11 | 51.8 | |
| OPSD+ViCuRModel Scale=2B2026.06 | 51.6 | |
| Qwen2.5-VL-7B-InstructModel scale=7B-8B2025.10 | 51.3 | |
| InternVL-3-8B-InstructModel scale=7B-8B2025.10 | 51.2 | |
| GPT-4oModel Category=Proprietary Models2025.06 | 48.5 | |
| Qwen2.5-VL-3B + PGPOModel Size=3B, Optimization Strategy=PGPO2026.04 | 48.45 | |
| Qwen2.5-VL-7BModel Size=7B2026.04 | 48.15 | |
| Qwen2.5-VL-3B + GRPOModel Size=3B, Optimization Strategy=GRPO2026.04 | 47.82 | |
| Qwen2.5-VL-3B + VPPOModel Size=3B, Optimization Strategy=VPPO2026.04 | 47.75 | |
| Qwen2.5-VL-3B + PAPOModel Size=3B, Optimization Strategy=PAPO2026.04 | 47.3 | |
| InternVL3-8BActivation Replay=false2025.11 | 46.5 | |
| Qwen2.5-VL-3B + DAPOModel Size=3B, Optimization Strategy=DAPO2026.04 | 45.86 | |
| Gemini-2.0-ProModel Category=Proprietary Models2025.06 | 43.3 | |
| Qwen2.5-VL-3BActivation Replay=false2025.11 | 42.7 | |
| Qwen2.5-VL-3B-InstructModel scale=3B-4B2025.10 | 42.5 | |
| InternVL-2.5-8BModel scale=7B-8B2025.10 | 41.9 | |
| MaLoRAModel=Qwen3-VL-8B, Training examples=4.0k2025.10 | 41.42 | |
| InternVL-2.5-4BModel scale=3B-4B2025.10 | 40.7 | |
| Claude-3.7-SonnetModel Category=Proprietary Models2025.06 | 39.7 | |
| PLM-HoneyBee-1BModel scale=1B2025.10 | 39.3 | |
| MaLoRAModel=Qwen2.5-VL-7B, Training examples=4.0k2025.10 | 38.52 |