Visual Mathematical Reasoning on MathVision
92.7AccuracyM3-ACE
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| M3-ACEAnchor Agent=Gemini 3 pro, Stage=1st Refine Select2026.03 | 92.7 | — | — | — | — | |
| M3-ACEAnchor Agent=Gemini 3 pro, Stage=1st Reflect All2026.03 | 89.1 | — | — | — | — | |
| M3-ACEAnchor Agent=Gemini 3 pro, Stage=1st Regenerate with Summary2026.03 | 88.1 | — | — | — | — | |
| Gemini 3-Pro2026.02 | 87.27 | — | — | — | — | |
| M3-ACEAnchor Agent=GPT 5, Stage=1st Refine Select2026.03 | 85.5 | — | — | — | — | |
| Gemini 3 pro (CoT)Anchor Agent=Gemini 3 pro, Stage=CoT Infer2026.03 | 85 | — | — | — | — | |
| M3-ACEAnchor Agent=Gemini 2.5 pro, Stage=1st Refine Select2026.03 | 84.4 | — | — | — | — | |
| M3-ACEAnchor Agent=GPT 5, Stage=1st Reflect All2026.03 | 82.2 | — | — | — | — | |
| M3-ACEAnchor Agent=GPT 5, Stage=1st Regenerate with Summary2026.03 | 81.3 | — | — | — | — | |
| M3-ACEAnchor Agent=Gemini 2.5 pro, Stage=1st Reflect All2026.03 | 81.2 | — | — | — | — | |
| M3-ACEAnchor Agent=Gemini 2.5 pro, Stage=1st Regenerate with Summary2026.03 | 80.5 | — | — | — | — | |
| M3-ACEAnchor Agent=Claude-45, Stage=1st Refine Select2026.03 | 79.2 | — | — | — | — | |
| GPT 5Mode=1st Regenerate with Summary2026.03 | 78.5 | — | — | — | — | |
| GPT-5tier=High2026.02 | 78.06 | — | — | — | — | |
| S1-VL-32B-RLParameter Count=32B, Training Phase=RL2026.04 | 77.7 | — | — | — | — | |
| S1-VL-32B-RL2026.06 | 77.7 | — | — | — | — | |
| Gemini 2.5 proMode=1st Regenerate with Summary2026.03 | 77.6 | — | — | — | — | |
| S1-Omni-Image2026.06 | 76.71 | — | — | — | — | |
| M3-ACEAnchor Agent=Claude-45, Stage=1st Reflect All2026.03 | 76.5 | — | — | — | — | |
| M3-ACEAnchor Agent=Claude-45, Stage=1st Regenerate with Summary2026.03 | 76 | — | — | — | — | |
| S1-VL-32B-SFTParameter Count=32B, Training Phase=SFT2026.04 | 75.89 | — | — | — | — | |
| S1-VL-32B-SFT2026.06 | 75.89 | — | — | — | — | |
| GPT-52026.04 | 75.66 | — | — | — | — | |
| GPT-52026.06 | 75.66 | — | — | — | — | |
| Qwen3-VL-235B-A22B-ThinkingParameter Count=235B-A22B, Reasoning Mode=Thinking2026.04 | 74.87 | — | — | — | — | |
| Qwen3-VL-235B-A22B-Thinking2026.06 | 74.87 | — | — | — | — | |
| ERNIE 5.02026.02 | 74.34 | — | — | — | — | |
| Gemini 2.5-Pro2026.02 | 73.3 | — | — | — | — | |
| Gemini 2.5 proMode=CoT Infer2026.03 | 73.3 | — | — | — | — | |
| Gemini 2.5 pro (CoT)Anchor Agent=Gemini 2.5 pro, Stage=CoT Infer2026.03 | 73.3 | — | — | — | — | |
| Gemini 2.5 Pro2026.04 | 73.3 | — | — | — | — | |
| GPT 5Mode=CoT Infer2026.03 | 72 | — | — | — | — | |
| GPT 5 (CoT)Anchor Agent=GPT 5, Stage=CoT Infer2026.03 | 72 | — | — | — | — | |
| Qwen3-VLmode=Thinking2026.02 | 71.84 | — | — | — | — | |
| Qwen3-VL-32B-ThinkingParameter Count=32B, Reasoning Mode=Thinking2026.04 | 71.51 | — | — | — | — | |
| Qwen3-VL-32B-Thinking2026.06 | 71.51 | — | — | — | — | |
| Claude-45Mode=1st Regenerate with Summary2026.03 | 71.2 | — | — | — | — | |
| Gemini 2.5 Flash2026.04 | 70.1 | — | — | — | — | |
| CodePercept-32B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 69.96 | — | — | — | — | |
| M3-ACEAnchor Agent=Claude-45, Stage=1st Reflect Reject2026.03 | 68.7 | — | — | — | — | |
| Gemini2.5-ProLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 66.8 | — | — | — | — | |
| M3-ACEAnchor Agent=Claude-45, Stage=1st Refine Reject2026.03 | 66.7 | — | — | — | — | |
| CodePercept-8B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 66.45 | — | — | — | — | |
| CodePercept-4B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 64.71 | — | — | — | — | |
| gemini-2.5-pro-exp-03-252026.01 | 63.5 | — | — | — | — | |
| Intern-S1Parameter Count=235B+6B2026.04 | 63.03 | — | — | — | — | |
| Qwen3-VL-32B-InstructLLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 62.66 | — | — | — | — | |
| CodePercept-32B-S1LLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 62.27 | — | — | — | — | |
| Claude-45Mode=CoT Infer2026.03 | 61.2 | — | — | — | — | |
| Claude-45 (CoT)Anchor Agent=Claude-45, Stage=CoT Infer2026.03 | 61.2 | — | — | — | — | |
| Qwen3-VL-235A22B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 60.43 | — | — | — | — | |
| o1Training Strategy=Close-source2025.08 | 60.3 | — | — | — | — | |
| o1Model Access=Closed-source2026.06 | 60.3 | — | — | — | — | |
| GPT5-ThinkingLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 60.03 | — | — | — | — | |
| Qwen3-VL-4B-InstructLLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 59.8 | — | — | — | — | |
| Qwen3-VL-8B-InstructLLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 59.67 | — | — | — | — | |
| Claude-Opus 4.1-ThinkingLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 59.61 | — | — | — | — | |
| CodePercept-8B-S1LLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 59.31 | — | — | — | — | |
| Qwen3-VL-32B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 58.55 | — | — | — | — | |
| InternVL3.5-8B-MPOTraining=DF-GRPO2026.03 | 58.2 | — | — | — | — | |
| CodePercept-4B-S1LLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 57.63 | — | — | — | — | |
| M3-ACEAnchor Agent=Gemini 3 pro, Stage=1st Reflect Reject2026.03 | 56.9 | — | — | — | — | |
| InternVL3.5-8B-MPOTraining=PRIME2026.03 | 56.9 | — | — | — | — | |
| InternVL3.5-8B-MPOTraining=GRPO2026.03 | 56.8 | — | — | — | — | |
| VisRefModel=Qwen3-VL-8B, Configuration=Visual token selection with adaptive stopping2026.02 | 56.6 | — | — | — | — | |
| M3-ACEAnchor Agent=GPT 5, Stage=1st Reflect Reject2026.03 | 55.6 | — | — | — | — | |
| Qwen3-VL-8B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 54.37 | — | — | — | — | |
| TSRModel=Qwen3-VL-8B, Configuration=Textual Self-Reflection2026.02 | 54.3 | — | — | — | — | |
| Qwen3-VL-4B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 54.21 | — | — | — | — | |
| Qwen2.5-VL-72BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 54.14 | — | — | — | — | |
| KeyeVL1.5-8BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 54.11 | — | — | — | — | |
| DPEBackbone=Qwen3-VL-8B-Instruct, Iteration=32026.02 | 53.88 | — | — | — | 64.39 | |
| STModel=Qwen3-VL-8B, Configuration=Standard Thinking (Baseline)2026.02 | 53.8 | — | — | — | — | |
| GLM-4.1V-9BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 53.75 | — | — | — | — | |
| Qwen3-VL-30A3B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 53.59 | — | — | — | — | |
| InternVL3.5-8BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 53.32 | — | — | — | — | |
| MiniCPM-V-4.5LLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 53.15 | — | — | — | — | |
| Claude4-Sonnet2026.02 | 52.7 | — | — | — | 64.1 | |
| InternVL3.5-8B-MPOTraining=MPO2026.03 | 52.6 | — | — | — | — | |
| Intern-S1-miniParameter Count=8B2026.04 | 52.47 | — | — | — | — | |
| Intern-S1-8BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 51.67 | — | — | — | — | |
| Gemini-2.0 proTraining Strategy=Close-source2025.08 | 48.1 | — | — | — | — | |
| M3-ACEAnchor Agent=GPT 5, Stage=1st Refine Reject2026.03 | 47.1 | — | — | — | — | |
| GPT5-Mini2026.02 | 46.6 | — | — | — | 53.8 | |
| OVR-7BCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B, Reasoning Mode=Text-based2025.09 | 46.4 | — | — | — | 55.76 | |
| M3-ACEAnchor Agent=Gemini 3 pro, Stage=1st Refine Reject2026.03 | 46.1 | — | — | — | — | |
| VisRefModel=InternVL3.5-8B, Configuration=Visual token selection with adaptive stopping2026.02 | 44.6 | — | — | — | — | |
| AgoraRouting=Pool of 5 baseline VLMs2026.01 | 44.3 | — | — | — | — | |
| InternVL3-78BScale=78B2026.01 | 43.1 | — | — | — | — | |
| M3-ACEAnchor Agent=Gemini 2.5 pro, Stage=1st Reflect Reject2026.03 | 41.6 | — | — | — | — | |
| Qwen2.5-VL-32BTraining=DF-GSPO2026.03 | 41.6 | — | — | — | — | |
| gemini-2.0-flash2026.01 | 41.3 | — | — | — | — | |
| Claude-3.7-SonnetTraining Strategy=Close-source2025.08 | 41.3 | — | — | — | — | |
| Claude-3.7-SonnetModel Access=Closed-source2026.06 | 41.3 | — | — | — | — | |
| Gemini-2.0-FlashModel Access=Closed-source2026.06 | 41.3 | — | — | — | — | |
| R1-ShareVL-32B2026.03 | 40.3 | — | — | — | — | |
| TSRModel=InternVL3.5-8B, Configuration=Textual Self-Reflection2026.02 | 40.1 | — | — | — | — | |
| InternVL2.5-38BModel Access=Open-source, Model Scale=38B2026.06 | 40.1 | — | — | — | — | |
| Qwen2.5-VL-32B-Instruct + NoisyRolloutStudent Model=Qwen2.5-VL-32B-Instruct, Distillation/Training Method=NoisyRollout2026.03 | 39.82 | — | — | — | — | |
| qwen2.5vl-72b-instructScale=72B2026.01 | 39.3 | — | — | — | — |