Logical Reasoning on LogicVista
81.4AccuracyGemini3-pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini3-proModel Category=Closed-Source SOTA Models2026.05 | 81.4 | |
| Gemini-2.5-ProModel Category=Proprietary Model2026.07 | 73.8 | |
| Doubao-Seed-1.6Model Category=Closed-Source SOTA Models2026.05 | 72.5 | |
| Qwen3-VL-235B-ThinkingModel Category=Open-Source Large Baselines2026.05 | 72.3 | |
| GPT-5-20250807Model Category=Closed-Source SOTA Models2026.05 | 70 | |
| GPT5Model Category=Proprietary Model2026.07 | 70 | |
| GLM-4.5VModel Category=Closed-Source SOTA Models2026.05 | 62.4 | |
| OSTBase Model=Qwen3-VL-8B-Instruct, Sampling Ratio=Best-20%2026.05 | 61.5 | |
| Qwen3-VL-8B-InstructStrategy=Base2026.05 | 61.3 | |
| DEITABase Model=Qwen3-VL-8B-Instruct, Sampling Ratio=Top-20%2026.05 | 61.2 | |
| Qwen3-VL-8B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 60.9 | |
| Qwen3-VL-8B-Instruct + RandomSampling Ratio=20%2026.05 | 60.6 | |
| OSTBase Model=Qwen3-VL-4B-Instruct, Sampling Ratio=Best-20%2026.05 | 60.4 | |
| Qwen3-VL-4B-InstructStrategy=Base2026.05 | 60 | |
| Qwen3-VL-8B-Instruct + Full SFTSampling Ratio=100%2026.05 | 60 | |
| Qwen3-VL-4B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 59.7 | |
| DEITABase Model=Qwen3-VL-4B-Instruct, Sampling Ratio=Top-20%2026.05 | 59.6 | |
| Qwen3-VL-4B-Instruct + RandomSampling Ratio=20%2026.05 | 59.5 | |
| Qwen3-VL-30B A3B-InstructParameters=30B A3B, Type=Instruct2025.12 | 58.2 | |
| Qwen3-VL-4B-Instruct + Full SFTSampling Ratio=100%2026.05 | 58.2 | |
| InternVL-3.5Model Size=8B2026.03 | 57.3 | |
| Qwen3-VL + LOCUSSize=4B2026.06 | 56.7 | |
| Qwen3-VL-8B + Qwen3-8B / H-OPDStudent=Qwen3-VL-4B-Instruct2026.07 | 55.9 | |
| Qwen3-VLModel Size=8B2026.03 | 55.3 | |
| Qwen3-VL-8B2026.07 | 55.3 | |
| MiMo-VL + LOCUSSize=7B2026.06 | 54.8 | |
| Penguin-VLModel Size=8B2026.03 | 53.8 | |
| Step-GUI-8BParameters=8B2025.12 | 53.7 | |
| GPT-4oTool Use=false, Param Size=-2026.03 | 53.2 | |
| Qwen3-VLSize=4B2026.06 | 52.5 | |
| MiMo-VLSize=7B2026.06 | 52.5 | |
| Qwen3-VL 8B-InstructParameters=8B, Type=Instruct2025.12 | 52.3 | |
| Step-GUI-4BParameters=4B2025.12 | 52.1 | |
| LASERModel Category=Open-source Vision-Language Model2026.07 | 51.2 | |
| Kimi-vl-A3B-thinkingModel Category=Open-Source Large Baselines2026.05 | 51 | |
| VAPO-Thinker-7BModel Category=Open-source Vision-Language Model2026.07 | 50.9 | |
| PRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 49.7 | |
| VisionR1-7BModel Category=Open-source Vision-Language Model2026.07 | 49.7 | |
| ThymeTool Use=true, Param Size=7B2026.03 | 49 | |
| DeepEyesV2Tool Use=true, Param Size=7B2026.03 | 48.7 | |
| VRETool Use=false, Param Size=7B2026.03 | 48.7 | |
| Qwen2.5-VL-7B + DUPLBase Model=Qwen2.5-VL-7B, RL Algorithm=DUPL2025.10 | 48.7 | |
| V-ZeroBackbone=Qwen2.5-VL-7B-Instruct, Iteration=22026.01 | 48.6 | |
| Qwen2.5-VL-7B + PGPOModel Size=7B, Optimization Strategy=PGPO2026.04 | 47.93 | |
| Qwen2.5-VL-7B + DAPOModel Size=7B, Optimization Strategy=DAPO2026.04 | 47.85 | |
| Supervised GRPOBackbone=Qwen2.5-VL-7B-Instruct2026.01 | 47.8 | |
| InternVL3.5-2BThinking Mode=true2025.12 | 47.7 | |
| DeepEyesParam Size=7B2025.05 | 47.7 | |
| InternVL3.5Parameters=2B2026.03 | 47.7 | |
| InternVL3.5#Params=2B2026.03 | 47.7 | |
| Qwen2.5-VL-7B + VPPOModel Size=7B, Optimization Strategy=VPPO2026.04 | 47.52 | |
| NoisyRollout-7BModel Scale=7B, Rollout=82026.06 | 47.5 | |
| VPPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 47.5 | |
| Qwen2.5-VL-7B + NoisyRolloutBase Model=Qwen2.5-VL-7B, RL Algorithm=NoisyRollout2025.10 | 47.5 | |
| NoisyRollout-7BModel Size=7B2026.04 | 47.49 | |
| Qwen2.5-VL-7B + PAPOModel Size=7B, Optimization Strategy=PAPO2026.04 | 47.46 | |
| V-ZeroBackbone=Qwen2.5-VL-7B-Instruct, Iteration=12026.01 | 47.4 | |
| VLAA-Thinker-7BBase Model=VLAA-Thinker-7B2025.10 | 47.3 | |
| MM-Eureka-Qwen-7BBase Model=MM-Eureka-Qwen-7B2025.10 | 47.3 | |
| Qwen2.5-VL-7B-Instruct (Base)Backbone=Qwen2.5-VL-7B-Instruct2026.01 | 47.2 | |
| OpenVLThinker-7BBase Model=OpenVLThinker-7B2025.10 | 47.1 | |
| PAPO-7BBase Model=PAPO-7B2025.10 | 47.1 | |
| Qwen2.5-VL + LOCUSSize=7B2026.06 | 47.1 | |
| OSTBase Model=Qwen3-VL-2B-Instruct, Sampling Ratio=Best-20%2026.05 | 47 | |
| Qwen2.5-VL-7B + GRPOModel Size=7B, Optimization Strategy=GRPO2026.04 | 46.81 | |
| Qwen3-VL-2B-InstructStrategy=Base2026.05 | 46.8 | |
| DEITABase Model=Qwen3-VL-2B-Instruct, Sampling Ratio=Top-20%2026.05 | 46.7 | |
| AutoToolSize=7B2026.05 | 46.7 | |
| PAPO-DModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 46.7 | |
| Qwen3-VL-2B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 46.5 | |
| VL-Rethinker-7BModel Scale=7B, Rollout=82026.06 | 46.5 | |
| VL-Rethinker-7BModel Size=7B2026.04 | 46.47 | |
| MM-Eureka-7BModel Size=7B2026.04 | 46.3 | |
| MM-Eureka-7BModel Scale=7B, Rollout=82026.06 | 46.3 | |
| InternVL3-8BSize=8B2026.05 | 46 | |
| Qwen2.5-VL-7B + GRPOBase Model=Qwen2.5-VL-7B, RL Algorithm=GRPO2025.10 | 46 | |
| Qwen2.5-VL*Param Size=7B2025.05 | 45.9 | |
| Qwen2.5-VLTool Use=false, Param Size=7B2026.03 | 45.9 | |
| Qwen3-VL-2B-Instruct + RandomSampling Ratio=20%2026.05 | 45.9 | |
| R1-ShareVL-7BModel Scale=7B, Rollout=82026.06 | 45.8 | |
| PAPO-GModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 45.8 | |
| R1-ShareVL-7BModel Size=7B2026.04 | 45.76 | |
| Qwen3-VL-2B-Instruct + Full SFTSampling Ratio=100%2026.05 | 45.6 | |
| GRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 45.6 | |
| R1-Onevision-7BModel Category=Open-source Vision-Language Model2026.07 | 45.6 | |
| Qwen2.5-VL-7BSize=7B2026.05 | 45.5 | |
| R1-Onevision-7BBase Model=R1-Onevision-7B2025.10 | 45.5 | |
| Qwen2.5-VL-7BBase Model=Qwen2.5-VL-7B2025.10 | 45.5 | |
| DeepEyesSize=7B2026.05 | 45.3 | |
| OpenVLThinker-7BBackbone=Qwen2.5-VL-7B-Instruct2026.01 | 45.1 | |
| Qwen2.5-VL-3B + DUPLBase Model=Qwen2.5-VL-3B, RL Algorithm=DUPL2025.10 | 45 | |
| PRPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 44.9 | |
| CFPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 44.45 | |
| BAGEL#Params=7B+7B2026.03 | 44.3 | |
| Qwen2.5-VLParam Size=7B2025.05 | 44.1 | |
| Qwen2.5-VL-3B + PGPOModel Size=3B, Optimization Strategy=PGPO2026.04 | 43.2 | |
| Qwen2.5-VLSize=7B2026.06 | 43.1 | |
| Qwen2.5-VL-7BModel Size=7B2026.04 | 43.01 | |
| Qwen2.5-VL-3B + PAPOModel Size=3B, Optimization Strategy=PAPO2026.04 | 42.7 | |
| Qwen2.5-VL-7BModel Category=Open-source Vision-Language Model2026.07 | 42.6 |