Multimodal Reasoning on Downstream Overall
55.22BoN@8 AccuracyInternVL2.5-38B + EVPV-PRM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVL2.5-38B + EVPV-PRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 55.22 | 9.78 | |
| Gemini-2.0-FlashReranking Strategy=Best-of-82026.03 | 53.4 | — | |
| InternVL2.5-38B + VisualPRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 50.7 | 6.3 | |
| Claude-3.5-SonnetReranking Strategy=Best-of-82026.03 | 50.5 | — | |
| GPT-4oReranking Strategy=Best-of-82026.03 | 47.9 | — | |
| InternVL2.5-26B + EVPV-PRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 46.75 | 9.52 | |
| InternVL2.5-26B + VisualPRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 45.8 | 8.9 | |
| InternVL2.5-38BPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 45.44 | — | |
| InternVL2.5-8B + EVPV-PRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 41.67 | 8.83 | |
| InternVL2.5-8B + VisualPRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 41.4 | 8.4 | |
| InternVL2.5-26BPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 37.23 | — | |
| InternVL2.5-8BPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 32.84 | — |