Visual Mathematical Reasoning on WeMath
98.7AccuracyOpenAI-o1
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| OpenAI-o1Learning Protocol=Closed-source VLM2026.02 | 98.7 | — | — | — | — | |
| o1Model Access=Closed-source2026.06 | 98.7 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=16, Generation Time (Gen.)=688.9, Total training time per step=1322.8, Learning Protocol=RLVR2026.02 | 84 | — | — | — | — | |
| Octopus-8B (Ours)Rollout count (rollout.n)=8, Generation Time (Gen.)=344.7, Total training time per step=958.12026.02 | 84 | — | — | — | — | |
| Qwen3-VL-8B-ThinkingLearning Protocol=Open-source Reasoning VLM2026.02 | 83 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + DAPORollout count (rollout.n)=16, Learning Protocol=RLVR2026.02 | 82.4 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=16, Generation Time (Gen.)=679.1, Total training time per step=1428.4, Learning Protocol=RLVR2026.02 | 78.5 | — | — | — | — | |
| Gemini-2.5-Pro2026.01 | 78 | — | — | — | — | |
| MiMo-VL-7B-SFTLearning Protocol=Open-source Reasoning VLM2026.02 | 77.7 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=16, Generation Time (Gen.)=816.6, Total training time per step=1543.3, Learning Protocol=RLVR2026.02 | 76.9 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=8, Generation Time (Gen.)=364.9, Total training time per step=845.1, Learning Protocol=RLVR2026.02 | 76.2 | — | — | — | — | |
| MiMo-VL-7B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 76.1 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=8, Generation Time (Gen.)=410.9, Total training time per step=895.8, Learning Protocol=RLVR2026.02 | 75.8 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=8, Generation Time (Gen.)=361.7, Total training time per step=753.1, Learning Protocol=RLVR2026.02 | 74 | — | — | — | — | |
| Claude-3.7-SonnetLearning Protocol=Closed-source VLM2026.02 | 72.6 | — | — | — | — | |
| Claude-3.7-SonnetTraining Strategy=Close-source2025.08 | 72.6 | — | — | — | — | |
| Claude-3.7-SonnetModel Access=Closed-source2026.06 | 72.6 | — | — | — | — | |
| QvQ-72B-PreviewModel Access=Open-source, Model Scale=72B2026.06 | 72.4 | — | — | — | — | |
| Shuffle-R1-Qwen-7BTraining Strategy=Zero RL2025.08 | 72.3 | — | — | — | — | |
| Gemini-2.0-FlashModel Access=Closed-source2026.06 | 71.4 | — | — | — | — | |
| GPT-52026.01 | 71.1 | — | — | — | — | |
| NoisyRollout-7B-K12Training Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 70.8 | — | — | — | — | |
| MMR1-Math-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 70.7 | — | — | — | — | |
| Qwen3-VL-8B-InstructLearning Protocol=Base VLM2026.02 | 70.5 | — | — | — | — | |
| R1-ShareVLModel Access=Open-source, Model Scale=7B, Learning Strategy=Visual-focused RL Fine-tuning2026.06 | 70.17 | — | — | — | — | |
| VL-Rethinker-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 70.1 | — | — | — | — | |
| Active-ZeroModel Backbone=Qwen2.5-VL-7B-Instruct2026.02 | 69.6 | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + VEPOModel Access=Open-source, Model Scale=7B, Learning Strategy=VEPO2026.06 | 69.54 | — | — | — | — | |
| VPPO-7BModel Access=Open-source, Model Scale=7B, Learning Strategy=Visual-focused RL Fine-tuning2026.06 | 69.43 | — | — | — | — | |
| ThinkLite-VL-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 69.3 | — | — | — | — | |
| NoisyRollout-7BModel Access=Open-source, Model Scale=7B, Learning Strategy=Visual-focused RL Fine-tuning2026.06 | 69.2 | — | — | — | — | |
| InternVL2.5-38BModel Access=Open-source, Model Scale=38B2026.06 | 69.1 | — | — | — | — | |
| GPT-4oModel Access=Closed-source2026.06 | 69 | — | — | — | — | |
| GPT-4oLearning Protocol=Closed-source VLM2026.02 | 68.8 | — | — | — | — | |
| GLM-4.5V2026.01 | 68.8 | — | — | — | — | |
| GPT-4oTraining Strategy=Close-source2025.08 | 68.8 | — | — | — | — | |
| PAPO-DAPO-7BModel Access=Open-source, Model Scale=7B, Learning Strategy=Visual-focused RL Fine-tuning2026.06 | 68.69 | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + GRPO (Top 40% high entropy)Model Access=Open-source, Model Scale=7B, Learning Strategy=GRPO, Entropy Selection Strategy=Top 40% high entropy2026.06 | 68.45 | — | — | — | — | |
| VisPlayModel Backbone=Qwen2.5-VL-7B-Instruct2026.02 | 67.99 | — | — | — | — | |
| VLAA-Thinker-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 67.7 | — | — | — | — | |
| Qwen2.5-VL-72BModel Access=Open-source, Model Scale=72B2026.06 | 67.5 | — | — | — | — | |
| MM-Eureka-Qwen-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 67.4 | — | — | — | — | |
| OpenVLThinker-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 67.2 | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + GRPOModel Access=Open-source, Model Scale=7B, Learning Strategy=GRPO2026.06 | 67.18 | — | — | — | — | |
| CHARTOOL-7BParameters=7B2026.04 | 67.13 | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + GRPO (Top 20% high entropy)Model Access=Open-source, Model Scale=7B, Learning Strategy=GRPO, Entropy Selection Strategy=Top 20% high entropy2026.06 | 67.13 | — | — | — | — | |
| VisionZero-CLEVRModel Backbone=Qwen2.5-VL-7B-Instruct2026.02 | 66.9 | — | — | — | — | |
| Shuffle-R1-Qwen-3BTraining Strategy=Zero RL2025.08 | 66.5 | — | — | — | — | |
| InternVL2.5-78BModel Access=Open-source, Model Scale=78B2026.06 | 66.3 | — | — | — | — | |
| InternVL3.5-8B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 65.8 | — | — | — | — | |
| Doubao-1.5-Pro2026.01 | 65.7 | — | — | — | — | |
| VisionZero-RealWorldModel Backbone=Qwen2.5-VL-7B-Instruct2026.02 | 65.4 | — | — | — | — | |
| Qwen2.5-VL-32BModel Access=Open-source, Model Scale=32B2026.06 | 65.4 | — | — | — | — | |
| Qwen2.5-VL-7BParameters=7B2026.04 | 64.94 | — | — | — | — | |
| EvoLMMModel Backbone=Qwen2.5-VL-7B-Instruct2026.02 | 64.89 | — | — | — | — | |
| VisionZero-ChartModel Backbone=Qwen2.5-VL-7B-Instruct2026.02 | 64.66 | — | — | — | — | |
| Base ModelModel Backbone=Qwen2.5-VL-7B-Instruct2026.02 | 64.48 | — | — | — | — | |
| COGFLOWModel Size=7B2026.01 | 64.1 | — | — | — | — | |
| GLM-4.1VModel Size=9B2026.01 | 63.8 | — | — | — | — | |
| Qwen2.5-VL-7BTraining Strategy=Open-Source SFT, Evaluation Toolkit=vLLM with custom scripts2025.08 | 63.5 | — | — | — | — | |
| RL-Evol + Verifier (full VeriEvol)Training Stage=RL, Model Size=7B, Evolution Strategy=Evolved, Verifier Usage=Yes (HTV-Agent)2026.06 | 63.05 | — | — | — | — | |
| Qwen2.5-VL-3B-Instruct + VEPOBase Model Size=3B, Fine-tuning=VEPO2026.06 | 62.7 | — | — | — | — | |
| VPPO-3BBase Model Size=3B2026.06 | 62.64 | — | — | — | — | |
| PAPO-DAPO-3BBase Model Size=3B2026.06 | 62.3 | — | — | — | — | |
| Qwen2.5-VL-7B-InstructModel Access=Open-source, Model Scale=7B2026.06 | 62.13 | — | — | — | — | |
| NoisyRollout-3BBase Model Size=3B2026.06 | 61.78 | — | — | — | — | |
| R1-ShareVLBase Model Size=3B2026.06 | 61.26 | — | — | — | — | |
| Active-ZeroModel Backbone=Qwen2.5-VL-3B-Instruct2026.02 | 61.21 | — | — | — | — | |
| Qwen2.5-VL-3B-Instruct + GRPOBase Model Size=3B, Fine-tuning=GRPO2026.06 | 61.15 | — | — | — | — | |
| Keye-VLModel Size=8B2026.01 | 60.7 | — | — | — | — | |
| R1-OneVision-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 60.6 | — | — | — | — | |
| Gemini2.5-ProLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 60.29 | — | — | — | — | |
| RL-EvolTraining Stage=RL, Model Size=7B, Evolution Strategy=Evolved, Verifier Usage=No2026.06 | 60 | — | — | — | — | |
| R1-VL-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 59.8 | — | — | — | — | |
| Qwen2.5-VL-3B-Instruct + GRPO (Top 40% high entropy)Base Model Size=3B, Fine-tuning=GRPO, Constraint=Top 40% high entropy2026.06 | 59.77 | — | — | — | — | |
| Qwen2.5-VL-3B-Instruct + GRPO (Top 20% high entropy)Base Model Size=3B, Fine-tuning=GRPO, Constraint=Top 20% high entropy2026.06 | 59.66 | — | — | — | — | |
| VeriEvol-SFT (VeriEvol-SFT-Init)Training Stage=SFT, Model Size=7B, Evolution Strategy=Evolved, Verifier Usage=No2026.06 | 59.33 | — | — | — | — | |
| CHARTOOL-3BParameters=3B2026.04 | 58.45 | — | — | — | — | |
| RL-OriginTraining Stage=RL, Model Size=7B, Evolution Strategy=None, Verifier Usage=No2026.06 | 58.19 | — | — | — | — | |
| VisPlayModel Backbone=Qwen2.5-VL-3B-Instruct2026.02 | 58.16 | — | — | — | — | |
| CodePercept-32B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 57.81 | — | — | — | — | |
| InternVL3.5Model Size=8B2026.01 | 57 | — | — | — | — | |
| Seed-only SFTTraining Stage=SFT, Model Size=7B, Evolution Strategy=None, Verifier Usage=No2026.06 | 56.76 | — | — | — | — | |
| Qwen3-VL-32B-InstructLLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 56.48 | — | — | — | — | |
| Qwen3-VL-8BTraining Strategy=Staged2026.05 | 56.1 | — | — | — | — | |
| Qwen3-VL-8BTraining=Staged2026.05 | 56.1 | — | — | — | — | |
| Qwen3-VL-8BTraining Strategy=Merged2026.05 | 55.43 | — | — | — | — | |
| GPT5-ThinkingLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 54.57 | — | — | — | — | |
| OneThinker-8BTraining=Base2026.05 | 54.57 | — | — | — | — | |
| CodePercept-32B-S1LLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 54.19 | — | — | — | — | |
| Qwen2.5-VL-3B-InstructBase Model Size=3B2026.06 | 53.97 | — | — | — | — | |
| OpenMMReasoner-7BTraining Stage=Baseline, Model Size=7B2026.06 | 53.81 | — | — | — | — | |
| Base ModelModel Backbone=Qwen2.5-VL-3B-Instruct2026.02 | 53.51 | — | — | — | — | |
| Qwen3-VL-235A22B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 53.05 | — | — | — | — | |
| CodePercept-8B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 52.29 | — | — | — | — | |
| Qwen2.5-VL-3BTraining Strategy=Open-Source SFT, Evaluation Toolkit=vLLM with custom scripts2025.08 | 51.7 | — | — | — | — | |
| Qwen3-VL-8BTraining Strategy=Base2026.05 | 50.86 | — | — | — | — | |
| Qwen3-VL-8BTraining=Base2026.05 | 50.86 | — | — | — | — | |
| PRCO-7BBackbone=Qwen2.5-VL-7B2026.03 | 50.29 | — | — | — | — | |
| CodePercept-4B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 50 | — | — | — | — |