Mathematical Reasoning on WeMath
84AccuracyOctopus
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| OctopusParameters=8B2026.07 | 84 | — | — | |
| Gemini 2.5 ProAccess=Close-source2026.02 | 80.6 | — | — | |
| Qwen3-VL + LOCUSSize=4B2026.06 | 78.8 | — | — | |
| MiMo-VL + LOCUSSize=7B2026.06 | 78.3 | — | — | |
| Gemini-2.5-Pro*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 78 | — | — | |
| Gemini-2.5-Pro*Model Category=Closed-Source Models2026.04 | 78 | 82.7 | — | |
| Gemini-2.5-ProModel Category=Proprietary Model2026.07 | 78 | — | — | |
| MiMo-VLSize=7B2026.06 | 77.2 | — | — | |
| Qwen3-VLSize=4B2026.06 | 76.3 | — | — | |
| Qwen2.5-VL-32B-Instruct + NoisyRolloutStudent Model=Qwen2.5-VL-32B-Instruct, Distillation/Training Method=NoisyRollout2026.03 | 75.51 | — | — | |
| RLR³Training Source=OpenMMR2026.05 | 74.3 | — | — | |
| Vision-R1Parameters=7B2026.07 | 73.9 | — | — | |
| MiMo-VLtraining=7B-SFT, parameters=7B2026.01 | 73.79 | — | — | |
| RLR³Training Source=DeepVision2026.05 | 73.6 | — | — | |
| Claude-3.5-Sonnet2026.02 | 73.05 | — | — | |
| RLVRTraining Source=DeepVision2026.05 | 72.9 | — | — | |
| EASE-7BModel Scale=7B2026.05 | 72.9 | — | — | |
| Claude 3.7Tool Use=false, Param Size=-2026.03 | 72.6 | — | — | |
| VGPO-7BModel Scale=7B2026.05 | 72.5 | — | — | |
| SRPOParameters=7B2026.07 | 71.6 | — | — | |
| RuCL2026.02 | 71.49 | — | — | |
| Semantic-backTool Use=false, Param Size=7B2026.03 | 71.3 | — | — | |
| BUS-8BParameters=8B2026.07 | 71.3 | — | — | |
| GPT-5-Thinking*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 71.1 | — | — | |
| GPT-5-Thinking*Model Category=Closed-Source Models2026.04 | 71.1 | 78.4 | — | |
| Qwen2.5-VL-7B + DUPLBase Model=Qwen2.5-VL-7B, RL Algorithm=DUPL2025.10 | 71.1 | — | — | |
| GPT5Model Category=Proprietary Model2026.07 | 71.1 | — | — | |
| Innovator-VLvariant=8B-Thinking, parameters=8B2026.01 | 70.86 | — | — | |
| PRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 70.8 | — | — | |
| Solution-backParameters=7B2026.07 | 70.8 | — | — | |
| MaLoRAModel=Qwen3-VL-8B, Training examples=5.8k2025.10 | 70.75 | — | — | |
| VPPO-RL-7BModel Scale=7B2026.05 | 70.6 | — | — | |
| NoisyRollout-7BParameters=7B2026.02 | 70.57 | — | — | |
| NoisyRolloutTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-7B-Instruct, Distillation/Training Method=NoisyRollout2026.03 | 70.57 | — | — | |
| RLR³Training Source=ViRL2026.05 | 70.4 | — | — | |
| GPT-5 miniVersion=high2026.05 | 70.2 | — | — | |
| VLAA-ThinkerParameters=7B2026.07 | 70.2 | — | — | |
| RKLTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-7B-Instruct, Distillation/Training Method=RKL2026.03 | 70.06 | — | — | |
| Official thinkingSource=Qwen3-VL2026.05 | 70 | — | — | |
| REOPOLDTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-7B-Instruct, Distillation/Training Method=REOPOLD2026.03 | 69.77 | — | — | |
| NoisyRollout-7BModel Scale=7B, Rollout=82026.06 | 69.6 | — | — | |
| Qwen2.5-VL-7B + NoisyRolloutBase Model=Qwen2.5-VL-7B, RL Algorithm=NoisyRollout2025.10 | 69.6 | — | — | |
| GeoFocusModel Scale=7B2026.02 | 69.4 | — | — | |
| PAPO_D-7BModel Scale=7B2026.05 | 69.4 | — | — | |
| VL-Rethinker-7BModel Scale=7B, Rollout=82026.06 | 69.3 | — | — | |
| Qwen2.5-VL-32BParameters=32B2026.02 | 69.1 | — | — | |
| VL-RethinkerParameters=7B2026.07 | 69.1 | — | — | |
| GPT4oAccess=Close-source2026.02 | 69 | — | — | |
| GPT-4oTool Use=false, Param Size=-2026.03 | 69 | — | — | |
| GRPOModel Scale=7B2026.02 | 68.9 | — | — | |
| Qwen2.5-VL-7B-Instruct + Faithful-MR1Backbone model=Qwen2.5-VL-7B-Instruct, Training Data=19.2K2026.05 | 68.9 | — | — | |
| VPPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 68.9 | — | — | |
| GPT-4o2026.02 | 68.8 | — | — | |
| VRETool Use=false, Param Size=7B2026.03 | 68.7 | — | — | |
| R1-ShareVL-7BModel Scale=7B, Rollout=82026.06 | 68.7 | — | — | |
| InternVL2.5-38BParameters=38B2026.02 | 68.61 | — | — | |
| Qwen2.5-VL-7B + GRPOBase Model=Qwen2.5-VL-7B, RL Algorithm=GRPO2025.10 | 68.5 | — | — | |
| Perception-R1-7BTraining Data=1.4K2026.05 | 68.3 | — | — | |
| RLVRTraining Source=ViRL2026.05 | 68.3 | — | — | |
| PAPO-DModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 68.3 | — | — | |
| CFPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 68.25 | — | — | |
| VL-Rethinker-7BParameters=7B2026.02 | 68.22 | — | — | |
| LoRAModel=Qwen3-VL-8B, Training examples=5.8k2025.10 | 68.1 | — | — | |
| GRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 68.1 | — | — | |
| VLAA-Thinker-7BBase Model=VLAA-Thinker-7B2025.10 | 67.7 | — | — | |
| PAPO-7BBase Model=PAPO-7B2025.10 | 67.6 | — | — | |
| VL-Rethinker-7BModel Scale=7B2026.05 | 67.5 | — | — | |
| GRPOTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-7B-Instruct, Distillation/Training Method=GRPO2026.03 | 67.4 | — | — | |
| RLVRTraining Source=OpenMMR2026.05 | 67.3 | — | — | |
| OpenVLThinker-7BBase Model=OpenVLThinker-7B2025.10 | 67.1 | — | — | |
| GRPOModel Scale=8B2026.06 | 67.05 | — | — | |
| MiMo-VLtraining=7B-RL, parameters=7B2026.01 | 67.01 | — | — | |
| PAPO-GModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 66.8 | — | — | |
| OpenVLThinkerParameters=7B2026.07 | 66.7 | — | — | |
| OpenVLThinker-7BParameters=7B2026.02 | 66.63 | — | — | |
| Semantic-back-7BParameters=7B2026.02 | 66.61 | — | — | |
| MM-Eureka-7BModel Scale=7B2026.05 | 66.6 | — | — | |
| MM-Eureka-7BModel Scale=7B, Rollout=82026.06 | 66.1 | — | — | |
| Qwen3-VLparameters=8B2026.01 | 66.03 | — | — | |
| Qwen3-VL-8BParameters=8B2026.07 | 66 | — | — | |
| Qwen2.5-VL-7B-Instruct + VPPOBackbone model=Qwen2.5-VL-7B-Instruct, Training Data=19.2K2026.05 | 65.7 | — | — | |
| MM-Eureka-Qwen-7BBase Model=MM-Eureka-Qwen-7B2025.10 | 65.6 | — | — | |
| Perception-R1-7BParameters=7B2026.02 | 65.57 | — | — | |
| ThinkLite-VL-7BParameters=7B2026.02 | 65.52 | — | — | |
| MM-Eureka-7BParameters=7B2026.02 | 65.34 | — | — | |
| Qwen2.5-VL + LOCUSSize=7B2026.06 | 65.1 | — | — | |
| Innovator-VLvariant=8B-Instruct, parameters=8B2026.01 | 65 | — | — | |
| Vision-R1-7BParameters=7B2026.02 | 64.98 | — | — | |
| SAYO-Qwen-8BBase Model=Qwen3-VL-8B2026.02 | 64.83 | — | — | |
| Qwen2.5-VL-7B-InstructStudent Model=Qwen2.5-VL-7B-Instruct2026.03 | 64.77 | — | — | |
| REOPOLDTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-3B-Instruct, Distillation/Training Method=REOPOLD2026.03 | 64.6 | — | — | |
| OpenVLThinker-7BParameters=7B2026.02 | 64.48 | — | — | |
| RKLTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-3B-Instruct, Distillation/Training Method=RKL2026.03 | 64.48 | — | — | |
| NoisyRollout-7BModel Scale=7B2026.05 | 64.4 | — | — | |
| Qwen2.5-VLSize=7B2026.06 | 64.4 | — | — | |
| Vision-SR1-7BTraining Data=56K2026.05 | 64.2 | — | — | |
| Qwen2.5-VL-3B + DUPLBase Model=Qwen2.5-VL-3B, RL Algorithm=DUPL2025.10 | 64.1 | — | — | |
| SAYO-Qwen-4BBase Model=Qwen3-VL-4B2026.02 | 63.97 | — | — | |
| MaLoRAModel=Qwen2.5-VL-7B, Training examples=5.8k2025.10 | 63.97 | — | — | |
| PRPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 63.9 | — | — |