Multimodal Reasoning on LogicVista
84.78AccuracyInternVL2.5-38B + VRPRM
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| InternVL2.5-38B + VRPRMBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 84.78 | — | — | — | — | |
| InternVL2.5-26B + VRPRMBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 83 | — | — | — | — | |
| InternVL2.5-8B + VRPRMBase Model=InternVL2.5-8B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 79.46 | — | — | — | — | |
| InternVL2.5-38B + VRPRM w/o RLBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 70.02 | — | — | — | — | |
| InternVL2.5-26B + VRPRM w/o RLBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 68.9 | — | — | — | — | |
| InternVL2.5-8B + VRPRM w/o RLBase Model=InternVL2.5-8B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 64.43 | — | — | — | — | |
| RTWIBase Model=Qwen3-VL Thinking, Setting=Online2026.02 | 61.7 | 50.5 | — | — | — | |
| GPT-4.1VLM Category=Proprietary VLMs2025.09 | 61.1 | — | — | — | — | |
| Athena-PRMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 60.9 | — | — | — | — | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 60.4 | — | — | — | — | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 60.4 | — | — | — | — | |
| Claude-3.5-SonnetSelection Strategy=Standard2025.08 | 60.4 | — | — | — | — | |
| Athena-ORMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 60.1 | — | — | — | — | |
| Self-consistencyPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 59.6 | — | — | — | — | |
| Qwen2.5-VL-72B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 59.1 | — | — | — | — | |
| RTWIBase Model=Qwen3-VL Instruct, Setting=Online2026.02 | 58.5 | 34.7 | — | — | — | |
| InternVL3-38BModel Category=Open-Source2025.06 | 58.4 | — | — | — | — | |
| Claude-3.7-SonnetModel Category=Proprietary2025.06 | 58.2 | — | — | — | — | |
| MM-Eureka-32BModel Category=Open-Source2025.06 | 58.2 | — | — | — | — | |
| QVQ-72B-PreviewModel Category=Open-Source2025.06 | 58.2 | — | — | — | — | |
| Claude3.7-SonnetVLM Category=Proprietary VLMs2025.09 | 58.2 | — | — | — | — | |
| Qwen2.5-VL-7B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 57.7 | — | — | — | — | |
| VL-Rethinker-72BModel Category=Open-Source2025.06 | 56.6 | — | — | — | — | |
| Qwen2.5-VL-3B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 56.4 | — | — | — | — | |
| OmniCaptionerMethod Category=Caption-then-Reason2025.06 | 56.2 | — | — | — | — | |
| DeepconfBase Model=Qwen3-VL Instruct, Setting=Online2026.02 | 56 | 32.1 | — | — | — | |
| InternVL3-78BModel Category=Open-Source2025.06 | 55.9 | — | — | — | — | |
| Qwen2.5-VL-72BModel Category=Open-Source2025.06 | 55.7 | — | — | — | — | |
| Qwen2.5-VL-72BReasoning Framework=Standard2025.06 | 55.7 | — | — | — | — | |
| Qwen2.5-VL-72BPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 55.7 | — | — | — | — | |
| Qwen2.5-VL-32BReasoning Framework=Standard2025.06 | 55 | — | — | — | — | |
| Self-Cer.Base Model=Qwen3-VL Instruct, Setting=Online2026.02 | 54.8 | 15.1 | — | — | — | |
| SCBase Model=Qwen3-VL Instruct, Setting=Online2026.02 | 54.6 | — | — | — | — | |
| ESCBase Model=Qwen3-VL Instruct, Setting=Online2026.02 | 54.6 | 7.9 | — | — | — | |
| ASCBase Model=Qwen3-VL Instruct, Setting=Online2026.02 | 54.1 | 27.6 | — | — | — | |
| LUSPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=LUSPO2026.02 | 53.7 | — | — | — | — | |
| InternVL2.5-38B + VisualPRMBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VisualPRM2025.08 | 53.7 | — | — | — | — | |
| CISCBase Model=Qwen3-VL Instruct, Setting=Online2026.02 | 53.4 | 14.1 | — | — | — | |
| Athena-PRMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 53 | — | — | — | — | |
| ARM-Thinker-7BParameters=7B2025.12 | 52.8 | — | — | — | — | |
| GPT-4o-20241120Model Category=Proprietary2025.06 | 52.8 | — | — | — | — | |
| GPT-4oSelection Strategy=Standard2025.08 | 52.8 | — | — | — | — | |
| Gemini-2.0-FlashModel Category=Proprietary2025.06 | 52.3 | — | — | — | — | |
| Gemini-2.0-FlashSelection Strategy=Standard2025.08 | 52.3 | — | — | — | — | |
| Revisual-R1Evaluation Source=original paper2026.05 | 52.3 | — | — | — | — | |
| DeepconfBase Model=Qwen3-VL Thinking, Setting=Online2026.02 | 51.8 | 34.5 | — | — | — | |
| AnE-3rdTraining Stage=Round 32026.05 | 51.3 | — | — | — | — | |
| Athena-ORMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 51.3 | — | — | — | — | |
| ReVisual-R1-7BModel Category=Open-Source2025.06 | 51.2 | — | — | — | — | |
| InternVL3-14BModel Category=Open-Source2025.06 | 51.2 | — | — | — | — | |
| InternVL2.5-26B + VisualPRMBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VisualPRM2025.08 | 51 | — | — | — | — | |
| SPARK-VL-7BEvaluation Source=original paper2026.05 | 50 | — | — | — | — | |
| OpenMMReasoner-7BEvaluation Source=original paper2026.05 | 50 | — | — | — | — | |
| Ovis2-34BModel Category=Open-Source2025.06 | 49.9 | — | — | — | — | |
| CISCBase Model=Qwen3-VL Thinking, Setting=Online2026.02 | 49.8 | 30.8 | — | — | — | |
| AnE-2ndTraining Stage=Round 22026.05 | 49.8 | — | — | — | — | |
| Metis-RISEEvaluation Source=original paper2026.05 | 49.7 | — | — | — | — | |
| PRCO-7BBackbone=Qwen2.5-VL-7B2026.03 | 49.66 | — | — | — | — | |
| Self-consistencyPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 49.5 | — | — | — | — | |
| MMR1-7B-RLBackbone=Qwen2.5-VL-7B2026.03 | 49.44 | — | — | — | — | |
| Self-Cer.Base Model=Qwen3-VL Thinking, Setting=Online2026.02 | 49.3 | 31.7 | — | — | — | |
| AnE-1stTraining Stage=Round 12026.05 | 49.1 | — | — | — | — | |
| ThymeTool=Code, Param Size=7B2025.11 | 49 | — | — | — | — | |
| Thyme-VL-7BVLM Category=Tool-Calling VLMs2025.09 | 49 | — | — | — | — | |
| MMR1Evaluation Source=original paper2026.05 | 48.8 | — | — | — | — | |
| R1-ShareVL-7BBackbone=Qwen2.5-VL-7B2026.03 | 48.76 | — | — | — | — | |
| DeepEyesV2Tool=General, Param Size=7B2025.11 | 48.7 | — | — | — | — | |
| ADHintEvaluation Source=original paper2026.05 | 48.7 | — | — | — | — | |
| Qwen2.5-VL-7B (Bo8)Training Strategy=GRPO, Sampling Strategy=Best-of-82025.06 | 48.6 | — | — | — | — | |
| PAPO-D-7BBackbone=Qwen2.5-VL-7B2026.03 | 48.54 | — | — | — | — | |
| Bagel-Zebra-CoT-7BVLM Category=Inner Visual Thought VLMs2025.09 | 48.4 | — | — | — | — | |
| NoisyRollout-7BBackbone=Qwen2.5-VL-7B2026.03 | 48.32 | — | — | — | — | |
| Vision-SR1-7BBackbone=Qwen2.5-VL-7B2026.03 | 48.32 | — | — | — | — | |
| MM-Eureka-7BModel Category=Open-Source2025.06 | 48.3 | — | — | — | — | |
| VisualPRM-8BPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 48.3 | — | — | — | — | |
| DeepSketcher-7BVLM Category=Inner Visual Thought VLMs2025.09 | 48.1 | — | — | — | — | |
| Vision-Matters-7BBackbone=Qwen2.5-VL-7B2026.03 | 48.09 | — | — | — | — | |
| CFPOGBackbone=Qwen2.5-VL-7B, Hyperparameters=gamma=0.02 + No Ent2026.06 | 48.07 | — | — | — | — | |
| BaseBase Model=Qwen3-VL Instruct, Setting=Online2026.02 | 47.9 | — | — | — | — | |
| InternVL2.5-38BBase Model=InternVL2.5-38B, Selection Strategy=Single-shot2025.08 | 47.9 | — | — | — | — | |
| Qwen2.5-VL-7BPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 47.9 | — | — | — | — | |
| DAPOBackbone=Qwen2.5-VL-7B2026.03 | 47.87 | — | — | — | — | |
| VPPO-7BBackbone=Qwen2.5-VL-7B2026.03 | 47.87 | — | — | — | — | |
| GSPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=GSPO2026.02 | 47.7 | — | — | — | — | |
| DeepEyesTool=Crop, Param Size=7B2025.11 | 47.7 | — | — | — | — | |
| DeepEyes-7BMethod Category=Tool-enabled2025.06 | 47.7 | — | — | — | — | |
| DeepEyes-7BVLM Category=Tool-Calling VLMs2025.09 | 47.7 | — | — | — | — | |
| DeepEyesTool=Yes2026.06 | 47.7 | — | — | — | — | |
| AIRTool=Yes2026.06 | 47.7 | — | — | — | — | |
| Ovis2-16BModel Category=Open-Source2025.06 | 47.4 | — | — | — | — | |
| Gemma-3-27BParameters=27B2025.12 | 47.3 | — | — | — | — | |
| Gemma-3-27BModel Category=Open-Source2025.06 | 47.3 | — | — | — | — | |
| PAPO-D-3BBackbone=Qwen2.5-VL-3B2026.03 | 47.2 | — | — | — | — | |
| Qwen2.5VL-7B w/ GRPOTool=No2026.06 | 47.2 | — | — | — | — | |
| RLVRBackbone=Qwen2.5VL-7B-Instruct, Averaged Over Seeds=true2026.07 | 46.7 | — | — | 52.2 | 0.495 | |
| RLVRBackbone=Qwen2.5VL-7B-Instruct, Evaluation Protocol=Averaged over 3 seeds (0, 42, 2025)2026.07 | 46.7 | — | — | 52.2 | 0.495 | |
| GRPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=GRPO2026.02 | 46.5 | — | — | — | — | |
| SCBase Model=Qwen3-VL Thinking, Setting=Online2026.02 | 46.4 | — | — | — | — | |
| ASCBase Model=Qwen3-VL Thinking, Setting=Online2026.02 | 46.4 | 21.2 | — | — | — | |
| MM-EurekaTool=✗, Param Size=7B2025.11 | 46.3 | — | — | — | — |