Visual Mathematical Reasoning on MathVista (BoN@8)
83.5BoN@8 AccuracyInternVL2.5-38B + EVPV-PRM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVL2.5-38B + EVPV-PRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 83.5 | 11.6 | |
| InternVL2.5-26B + EVPV-PRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 79.6 | 11.4 | |
| InternVL2.5-8B + EVPV-PRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 76.3 | 11.8 | |
| InternVL2.5-38B + VisualPRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 73.9 | 2 | |
| InternVL2.5-26B + VisualPRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 73.1 | 4.9 | |
| InternVL2.5-38BPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 71.9 | — | |
| Gemini-2.0-FlashReranking Strategy=Best-of-82026.03 | 70.4 | — | |
| InternVL2.5-8B + VisualPRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 68.5 | 4 | |
| InternVL2.5-26BPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 68.2 | — | |
| Claude-3.5-SonnetReranking Strategy=Best-of-82026.03 | 65.3 | — | |
| InternVL2.5-8BPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 64.5 | — | |
| PEPO_GBackbone=Qwen2.5-VL-3B-Instruct, Training Data=ViRL39K2026.03 | 63.48 | — | |
| PAPO_GBackbone=Qwen2.5-VL-3B-Instruct, Training Data=ViRL39K2026.03 | 61.38 | — | |
| GPT-4oReranking Strategy=Best-of-82026.03 | 60 | — | |
| GRPOBackbone=Qwen2.5-VL-3B-Instruct, Training Data=ViRL39K2026.03 | 59.34 | — |