Mathematical Reasoning on MathVerse
84.1AccuracyGPT-5-high
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5-highModel Category=Proprietary Models2026.06 | 84.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ERNIE 5.0-BaseModel type=pre-trained2026.02 | 81.45 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-5-Thinking*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 81.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT5Model Category=Proprietary Model2026.07 | 81.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-2.5-Pro*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 76.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-2.5-ProModel Category=Proprietary Model2026.07 | 76.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 72.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VPPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 71 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-Rethinker-7BModel Scale=7B, Rollout=82026.06 | 68.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PAPO-DModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 68.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OctopusParameters=8B2026.07 | 68.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PAPO-GModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 68.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| R1-ShareVL-7BModel Scale=7B, Rollout=82026.06 | 68.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-Opus-4.1Model Category=Proprietary Models2026.06 | 68.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NoisyRollout-7BModel Scale=7B, Rollout=82026.06 | 67.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM-Eureka-7BModel Scale=7B, Rollout=82026.06 | 67.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 66.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PRPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 62.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MathFlowBase LLM=Gemini 2.5-pro, Perception Model=MathFlow-P-7B2025.03 | 62.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| AnEParameters=7B2026.07 | 62.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VLModel Size=8B2026.03 | 62.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CFPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 61.92 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OPD+ViCuRModel Scale=8B2026.06 | 61.55 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VLSize=4B2026.06 | 61.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VPPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 61.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL + LOCUSSize=4B2026.06 | 61.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MathFlowBase LLM=Claude-sonnet-3.5, Perception Model=MathFlow-P-7B2025.03 | 60.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PAPO-DModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 60.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TVI-CoTModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=8B2026.06 | 60.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OPDModel Scale=8B2026.06 | 60.05 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini 2.5-pro2025.03 | 59.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-3.5-Sonnet2026.02 | 57.64 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-32BParameters=32B2026.02 | 57.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PAPO-GModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 57.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-sonnet-3.52025.03 | 57.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MiMo-VL + LOCUSSize=7B2026.06 | 57 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MathFlowBase LLM=InternVL-2.5-78B, Perception Model=MathFlow-P-7B2025.03 | 56.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MathFlowBase LLM=GPT-4V, Perception Model=MathFlow-P-7B2025.03 | 56.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B (Baseline)Model Category=Baseline, Model Scale=8B2026.06 | 56.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DAPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 56.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOModel Scale=8B2026.06 | 56.22 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BUS-8BParameters=8B2026.07 | 56.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL-3.5Model Size=8B2026.03 | 55.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SRPOParameters=7B2026.07 | 55.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DAPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 55.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 55.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LASERModel Category=Open-source Vision-Language Model2026.07 | 55.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-Cogito-7B + LEADBackbone=VL-Cogito-7B, Decoding Strategy=LEAD2026.03 | 55.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-Rethinker-7B + LEADBackbone=VL-Rethinker-7B, Decoding Strategy=LEAD2026.03 | 54.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-2.0-FlashSampling Strategy=avg@8, Inference Temperature=12026.01 | 54.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| V-STARData Size=40k2026.04 | 54.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vision-R1-7B + LEADBackbone=Vision-R1-7B, Decoding Strategy=LEAD2026.03 | 54.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4V2025.03 | 54.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-Rethinker-7BBackbone=VL-Rethinker-7B2026.03 | 54.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-RethinkerData Size=39k2026.04 | 54.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-Rethinker-7BModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=7B2026.06 | 54.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MiMo-VLSize=7B2026.06 | 54.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RuCL2026.02 | 54.14 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-Rethinker-7BParameters=7B2026.02 | 53.86 | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5#Params=2B2026.03 | 53.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Base ModelModel Scale=8B2026.06 | 53.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-Cogito-7BBackbone=VL-Cogito-7B2026.03 | 53.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VAPO-Thinker-7BModel Category=Our models, Parameter Scale=7B2025.09 | 53.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-CogitoData Size=80k2026.04 | 53.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VAPO-Thinker-7BModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=7B2026.06 | 53.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VAPO-Thinker-7BModel Category=Open-source Vision-Language Model2026.07 | 53.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OPSD+ViCuRModel Scale=8B2026.06 | 53.25 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vision-R1-7BParameters=7B2026.02 | 53.23 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VRETool Use=false, Param Size=7B2026.03 | 53.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OPSDModel Scale=8B2026.06 | 52.92 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-RethinkerParameters=7B2026.07 | 52.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepEyesV2Tool Use=true, Param Size=7B2026.03 | 52.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VLAA-ThinkerParameters=7B2026.07 | 52.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Perception-R1-7BParameters=7B2026.02 | 52.56 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vision-R1-7BBackbone=Vision-R1-7B2026.03 | 52.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vision-R1-7BModel Category=Open-source, Parameter Scale=7B2025.09 | 52.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vision-R1Data Size=210k2026.04 | 52.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vision-R1-7BModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=7B2026.06 | 52.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisionR1-7BModel Category=Open-source Vision-Language Model2026.07 | 52.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vision-R1Parameters=7B2026.07 | 52.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VLParameters=2B2026.03 | 52.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ThinkLite-VLData Size=11k2026.04 | 52.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B + DUPLBase Model=Qwen2.5-VL-7B, RL Algorithm=DUPL2025.10 | 52.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude 3.7Tool Use=false, Param Size=-2026.03 | 52 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PAPO-7BBase Model=PAPO-7B2025.10 | 52 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8BParameters=8B2026.07 | 52 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + Faithful-MR1Backbone model=Qwen2.5-VL-7B-Instruct, Training Data=19.2K2026.05 | 51.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PAPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 51.89 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Solution-backParameters=7B2026.07 | 51.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ThinkLite-VL-7BParameters=7B2026.02 | 51.47 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM-Eureka-7BParameters=7B2026.02 | 51.09 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B + NoisyRolloutBase Model=Qwen2.5-VL-7B, RL Algorithm=NoisyRollout2025.10 | 51 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Penguin-VLModel Size=8B2026.03 | 50.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oTool Use=false, Param Size=-2026.03 | 50.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Semantic-backTool Use=false, Param Size=7B2026.03 | 50.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OpenVLThinkerData Size=59.2k2026.04 | 50.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM-Eureka-Qwen-7BBase Model=MM-Eureka-Qwen-7B2025.10 | 50.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OpenVLThinker-7BModel Category=Open-source Vision-Language Model2026.07 | 50.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o2026.02 | 50.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + VPPOBackbone model=Qwen2.5-VL-7B-Instruct, Training Data=19.2K2026.05 | 49.5 | — | — | — | — | — | — | — | — | — | — | — | — |