Multimodal Reasoning on MathVista (Accuracy)
83.5AccuracyInternVL2.5-38B + VRPRM
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVL2.5-38B + VRPRMBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 83.5 | |
| InternVL2.5-26B + VRPRMBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 81.2 | |
| AnE-3rdTraining Stage=Round 32026.05 | 81.2 | |
| AnE-2ndTraining Stage=Round 22026.05 | 80.1 | |
| AnE-1stTraining Stage=Round 12026.05 | 79.6 | |
| OpenMMReasoner-7BEvaluation Source=original paper2026.05 | 79.5 | |
| InternVL2.5-8B + VRPRMBase Model=InternVL2.5-8B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 79.1 | |
| Athena-PRMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 79.1 | |
| InternVL2.5-38B + VRPRM w/o RLBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 78.4 | |
| Preliminary RLTraining Phase=Warm-up2026.05 | 77.9 | |
| Athena-ORMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 77.8 | |
| InternVL2.5-26B + VRPRM w/o RLBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 77.4 | |
| Self-consistencyPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 77 | |
| SPARK-VL-7BEvaluation Source=original paper2026.05 | 75.9 | |
| Metis-RISEEvaluation Source=original paper2026.05 | 75.8 | |
| SRPOEvaluation Source=original paper2026.05 | 75.8 | |
| Athena-PRMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 75.2 | |
| VL-RethinkerEvaluation Source=original paper2026.05 | 74.9 | |
| ADHintEvaluation Source=original paper2026.05 | 74.4 | |
| Qwen2.5-VL-72BPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 74.2 | |
| LLaVA-Critic-R1Evaluation Source=original paper2026.05 | 74 | |
| InternVL2.5-38B + VisualPRMBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VisualPRM2025.08 | 73.9 | |
| Vision-R1Evaluation Source=reproduced2026.05 | 73.5 | |
| InternVL2.5-26B + VisualPRMBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VisualPRM2025.08 | 73.1 | |
| Revisual-R1Evaluation Source=original paper2026.05 | 73.1 | |
| Athena-ORMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 72.8 | |
| InternVL2.5-8B + VRPRM w/o RLBase Model=InternVL2.5-8B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 72.6 | |
| OpenVLThinkerEvaluation Source=original paper2026.05 | 72.3 | |
| MMR1Evaluation Source=original paper2026.05 | 72 | |
| InternVL2.5-38BBase Model=InternVL2.5-38B, Selection Strategy=Single-shot2025.08 | 71.9 | |
| Self-consistencyPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 71.6 | |
| Athena-PRMPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 71.4 | |
| Gemini-2.0-FlashSelection Strategy=Standard2025.08 | 70.4 | |
| VisualPRM-8BPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 70.3 | |
| C2-EvoEvaluation Source=original paper2026.05 | 68.7 | |
| InternVL2.5-8B + VisualPRMBase Model=InternVL2.5-8B, Selection Strategy=Best-of-8, Critic Model=VisualPRM2025.08 | 68.5 | |
| VisualPRM-8BPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 68.5 | |
| InternVL2.5-26BBase Model=InternVL2.5-26B, Selection Strategy=Single-shot2025.08 | 68.2 | |
| Qwen2.5-VL-7B-InstructEvaluation Source=original paper2026.05 | 68.2 | |
| Qwen2.5-VL-7BPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 68.1 | |
| Athena-ORMPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 66.9 | |
| Self-consistencyPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 66.1 | |
| Claude-3.5-SonnetSelection Strategy=Standard2025.08 | 65.3 | |
| InternVL2.5-8BBase Model=InternVL2.5-8B, Selection Strategy=Single-shot2025.08 | 64.5 | |
| InternVL2.5-8BPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 64.5 | |
| GPT-4oSelection Strategy=Standard2025.08 | 60 | |
| BaselineModel=Qwen2 [92]2026.07 | 56.1 | |
| ESCModel=Qwen2 [92]2026.07 | 56.1 | |
| ESCModel=LLaVA [50]2026.07 | 23.4 | |
| BaselineModel=LLaVA [50]2026.07 | 22.2 |