Visual Mathematical Reasoning on MathVision (BoN@8)
43.6BoN@8 AccuracyGemini-2.0-Flash
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini-2.0-FlashReranking Strategy=Best-of-82026.03 | 43.6 | — | |
| InternVL2.5-38B + EVPV-PRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 37.59 | 5.39 | |
| Claude-3.5-SonnetReranking Strategy=Best-of-82026.03 | 35.6 | — | |
| InternVL2.5-38B + VisualPRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 35.2 | 3 | |
| InternVL2.5-38BPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 32.2 | — | |
| GPT-4oReranking Strategy=Best-of-82026.03 | 31.2 | — | |
| InternVL2.5-26B + VisualPRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 29.6 | 6.2 | |
| InternVL2.5-26B + EVPV-PRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 28.11 | 4.71 | |
| InternVL2.5-8B + VisualPRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 25.7 | 8.7 | |
| InternVL2.5-26BPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 23.4 | — | |
| InternVL2.5-8B + EVPV-PRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 22.07 | 5.07 | |
| InternVL2.5-8BPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 17 | — |