Multimodal Reasoning on MathVision
59.41AccuracyInternVL2.5-38B + VRPRM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVL2.5-38B + VRPRMBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 59.41 | — | |
| MiMo-VL-7BBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=No2026.03 | 57.9 | — | |
| MiMo-VL-7B +PCBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 57 | — | |
| InternVL2.5-26B + VRPRMBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 55.79 | — | |
| AutoNPO2026.04 | 55.72 | — | |
| ExGRPOtype=historical replay2026.04 | 55.46 | — | |
| NPOstage=early + late-stage2026.04 | 54.61 | — | |
| NPOstage=early-stage only2026.04 | 54.31 | — | |
| RLEPtype=far future2026.04 | 54.23 | — | |
| LUFFYtype=external teacher2026.04 | 54 | — | |
| Qwen2.5-VL-72B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 53.4 | — | |
| RLSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 52.73 | — | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 52.1 | — | |
| InternVL2.5-8B + VRPRMBase Model=InternVL2.5-8B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 51.44 | — | |
| GRPOBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 48.82 | — | |
| GRPOtype=pure on-policy2026.04 | 48.82 | — | |
| Revisual-R1Evaluation Source=original paper2026.05 | 48.8 | — | |
| GRPO + OPSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 48.52 | — | |
| OPSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 47.53 | — | |
| Base LLMBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 47.37 | — | |
| Qwen3-VL-8B-Instruct2026.04 | 47.37 | — | |
| SDPOBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 47.27 | — | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 47 | — | |
| GPT-4.1VLM Category=Proprietary VLMs2025.09 | 46.4 | — | |
| Athena-PRMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 44.8 | — | |
| AnE-3rdTraining Stage=Round 32026.05 | 43.9 | — | |
| Qwen2.5-VL-7B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 43.7 | — | |
| Gemini-2.0-FlashModel Category=Proprietary2025.06 | 43.6 | — | |
| Gemini-2.0-FlashSelection Strategy=Standard2025.08 | 43.6 | — | |
| OpenMMReasoner-7BEvaluation Source=original paper2026.05 | 43.6 | — | |
| InternVL2.5-38B + VRPRM w/o RLBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 43.45 | — | |
| OmniCaptionerMethod Category=Caption-then-Reason2025.06 | 43.3 | — | |
| InternVL3-78BModel Category=Open-Source2025.06 | 43.1 | — | |
| Self-consistencyPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 43.1 | — | |
| ReVisual-R1-7BModel Category=Open-Source2025.06 | 43 | — | |
| Athena-ORMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 43 | — | |
| ECSOMethod Category=Caption-then-Reason2025.06 | 42.7 | — | |
| VL-Rethinker-72BModel Category=Open-Source2025.06 | 42.2 | — | |
| Claude-3.7-SonnetModel Category=Proprietary2025.06 | 41.9 | — | |
| Claude3.7-SonnetVLM Category=Proprietary VLMs2025.09 | 41.9 | — | |
| AnE-2ndTraining Stage=Round 22026.05 | 40.9 | — | |
| Qwen2.5-VL-3B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 40.8 | — | |
| Gemma-3-27BModel Category=Open-Source2025.06 | 39.8 | — | |
| AnE-1stTraining Stage=Round 12026.05 | 39.7 | — | |
| Qwen2.5-VL-72BModel Category=Open-Source2025.06 | 39.3 | — | |
| Qwen2.5-VL-72BReasoning Framework=Standard2025.06 | 39.3 | — | |
| Qwen2.5-VL-72BPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 39.3 | — | |
| InternVL2.5-26B + VRPRM w/o RLBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 37.99 | — | |
| Qwen2.5-VL-32BReasoning Framework=Standard2025.06 | 37.8 | — | |
| InternVL3-14BModel Category=Open-Source2025.06 | 37.2 | — | |
| MM-Eureka-32BModel Category=Open-Source2025.06 | 36.6 | — | |
| Claude-3.5-SonnetSelection Strategy=Standard2025.08 | 35.6 | — | |
| InternVL2.5-38B + VisualPRMBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VisualPRM2025.08 | 35.2 | — | |
| QVQ-72B-PreviewModel Category=Open-Source2025.06 | 34.9 | — | |
| InternVL3-38BModel Category=Open-Source2025.06 | 34.2 | — | |
| InternVL2.5-8B + VRPRM w/o RLBase Model=InternVL2.5-8B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 33.95 | — | |
| SRPO-7BMethod Category=Verification-augmented2025.06 | 32.9 | — | |
| SRPOEvaluation Source=original paper2026.05 | 32.9 | — | |
| Preliminary RLTraining Phase=Warm-up2026.05 | 32.8 | — | |
| Athena-PRMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 32.5 | — | |
| DeepSketcher-7BVLM Category=Inner Visual Thought VLMs2025.09 | 32.3 | — | |
| VL-RethinkerEvaluation Source=original paper2026.05 | 32.3 | — | |
| InternVL2.5-38BBase Model=InternVL2.5-38B, Selection Strategy=Single-shot2025.08 | 32.2 | — | |
| Ovis2-34BModel Category=Open-Source2025.06 | 31.9 | — | |
| MMR1Evaluation Source=original paper2026.05 | 31.8 | — | |
| Qwen2.5-VL-7B (Bo8)Training Strategy=GRPO, Sampling Strategy=Best-of-82025.06 | 31.6 | — | |
| VisualPRM-8BPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 31.3 | — | |
| TGRL-DAPOData=39K, Base Model=Qwen-2.5-VL-7B2026.03 | 31.26 | — | |
| GPT-4o-20241120Model Category=Proprietary2025.06 | 31.2 | — | |
| GPT-4oSelection Strategy=Standard2025.08 | 31.2 | — | |
| SPARK-VL-7BEvaluation Source=original paper2026.05 | 31.1 | — | |
| NoisyRolloutData=6.4K2026.03 | 30.6 | — | |
| LLaVA-Critic-R1Evaluation Source=original paper2026.05 | 30.6 | — | |
| MixedR1 7BModel Category=Reasoning MLLMs2026.01 | 30.3 | — | |
| Gemma-3-12BModel Category=Open-Source2025.06 | 30.3 | — | |
| CoT Cold-Start + Solver FeedbackBackbone=Qwen2.5-VL-7B-Instruct2025.11 | 30.26 | — | |
| Ovis2-16BModel Category=Open-Source2025.06 | 30.1 | — | |
| VL-Rethinker-7BModel Category=Open-Source2025.06 | 30 | — | |
| R1-Onevision 7BModel Category=Reasoning MLLMs2026.01 | 29.9 | — | |
| VisonR1 7BModel Category=Reasoning MLLMs2026.01 | 29.9 | — | |
| TGRL-GRPOData=39K, Base Model=Qwen-2.5-VL-7B2026.03 | 29.87 | — | |
| Athena-ORMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 29.8 | — | |
| CodeDanceTooling Type=Dynamic Tooling2026.02 | 29.6 | — | |
| Self-InstructBackbone=Qwen2.5-VL-7B-Instruct2025.11 | 29.6 | — | |
| InternVL2.5-26B + VisualPRMBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VisualPRM2025.08 | 29.6 | — | |
| InternVL3Tool=✗, Param Size=8B2025.11 | 29.3 | — | |
| InternVL3-8BModel Category=Open-Source2025.06 | 29.3 | — | |
| Seed SetBackbone=Qwen2.5-VL-7B-Instruct2025.11 | 29.28 | — | |
| DAPOData=39K, Base Model=Qwen-2.5-VL-7B2026.03 | 28.94 | — | |
| DeepEyes-v2Tooling Type=Dynamic Tooling2026.02 | 28.9 | — | |
| DeepEyesV2Tool=General, Param Size=7B2025.11 | 28.9 | — | |
| GRPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=GRPO2026.02 | 28.7 | — | |
| PyVision-ImageTooling Type=Dynamic Tooling2026.02 | 28.7 | — | |
| Metis-RISEEvaluation Source=original paper2026.05 | 28.7 | — | |
| TGRL-DAPOData=2.1K, Base Model=Qwen-2.5-VL-7B2026.03 | 28.61 | — | |
| Perception-R1Data=1.4K2026.03 | 28.6 | — | |
| Mirage-7BVLM Category=Inner Visual Thought VLMs2025.09 | 28.6 | — | |
| Self-consistencyPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 28.6 | — | |
| VL-RethinkerTool=✗, Param Size=7B2025.11 | 28.4 | — | |
| GRPOData=39K, Base Model=Qwen-2.5-VL-7B2026.03 | 28.32 | — |