Visual Logical Reasoning on LogicVista (BoN@8)
60.4BoN@8 AccuracyClaude-3.5-Sonnet
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Claude-3.5-SonnetReranking Strategy=Best-of-82026.03 | 60.4 | — | |
| InternVL2.5-38B + EVPV-PRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 58.74 | 10.84 | |
| InternVL2.5-38B + VisualPRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 53.7 | 5.8 | |
| GPT-4oReranking Strategy=Best-of-82026.03 | 52.8 | — | |
| Gemini-2.0-FlashReranking Strategy=Best-of-82026.03 | 52.3 | — | |
| InternVL2.5-26B + EVPV-PRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 51.72 | 12.08 | |
| InternVL2.5-26B + VisualPRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 51 | 11.4 | |
| InternVL2.5-38BPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 47.9 | — | |
| InternVL2.5-8B + EVPV-PRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 45.33 | 8.95 | |
| InternVL2.5-8B + VisualPRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 43.8 | 7.8 | |
| InternVL2.5-26BPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 39.64 | — | |
| InternVL2.5-8BPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 36.38 | — |