Mathematical Reasoning on LogicVista
73.8AccuracyGemini-2.5-Pro*
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini-2.5-Pro*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 73.8 | — | — | |
| GPT-5-Thinking*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 70 | — | — | |
| GPT-4oActivation Replay=false2025.11 | 64.4 | — | — | |
| MM-Eureka-Qwen-32BActivation Replay=true2025.11 | 54.6 | — | — | |
| MM-Eureka-Qwen-32BActivation Replay=false2025.11 | 52.8 | — | — | |
| Gemini-2.0-FlashActivation Replay=false2025.11 | 52.3 | — | — | |
| MM-Eureka-Qwen-7BActivation Replay=true2025.11 | 51 | — | — | |
| VAPO-Thinker-7BModel Category=Our models, Parameter Scale=7B2025.09 | 50.9 | — | — | |
| Vision-R1-7BModel Category=Open-source, Parameter Scale=7B2025.09 | 49.7 | — | — | |
| VL-Rethinker-7BActivation Replay=true2025.11 | 49.7 | — | — | |
| Claude-3.5 SonnetActivation Replay=false2025.11 | 49.3 | — | — | |
| MM-Eureka-Qwen-7BActivation Replay=false2025.11 | 49.2 | — | — | |
| VLAA-Thinker-7BModel Category=Open-source, Parameter Scale=7B2025.09 | 48.5 | — | — | |
| VL-Rethinker-7BActivation Replay=false2025.11 | 46.1 | — | — | |
| R1-OneVision-7BModel Category=Open-source, Parameter Scale=7B2025.09 | 45.6 | — | — | |
| MMR1-Math-v0-7BActivation Replay=true2025.11 | 44.5 | — | — | |
| InternVL3-8BActivation Replay=false2025.11 | 43.6 | — | — | |
| MMR1-Math-v0-7BActivation Replay=false2025.11 | 43.6 | — | — | |
| Kimi-VL-16BActivation Replay=false2025.11 | 42.7 | — | — | |
| QvQ-72B-PreviewActivation Replay=false2025.11 | 42.7 | — | — | |
| Qwen2.5-VL-7BModel Category=Open-source, Parameter Scale=7B2025.09 | 42.6 | — | — | |
| GPT-4o-miniActivation Replay=false2025.11 | 41.4 | — | — | |
| VAPO-Thinker-3BModel Category=Our models, Parameter Scale=3B2025.09 | 39.7 | — | — | |
| InternVL2.5-8B*Model Category=Open-source, Parameter Scale=8B, Reference Source=OpenCompass leaderboard2025.09 | 38.3 | — | — | |
| Qwen2.5-VL-7BActivation Replay=false2025.11 | 35.6 | — | — | |
| LLaVA-OV-7BActivation Replay=false2025.11 | 32 | — | — | |
| Qwen2.5-VL-3BActivation Replay=false2025.11 | 26.6 | — | — | |
| DAPOModel Scale=Qwen2.5-VL-3B2026.06 | — | 39.8 | — | |
| DAPOModel Scale=Qwen2.5-VL-7B2026.06 | — | 42.3 | — | |
| DAPO + DyCo-RLModel Scale=Qwen2.5-VL-3B2026.06 | — | 41.4 | — | |
| DAPO + DyCo-RLModel Scale=Qwen2.5-VL-7B2026.06 | — | 43.8 | — | |
| Gemini-2.5-Pro*Model Category=Closed-Source Models2026.04 | — | 73.8 | 76.4 | |
| GPT-5-Thinking*Model Category=Closed-Source Models2026.04 | — | 70 | 81.5 | |
| GRPOModel Scale=Qwen2.5-VL-3B2026.06 | — | 39.8 | — | |
| GRPOModel Scale=Qwen2.5-VL-7B2026.06 | — | 46.2 | — | |
| GRPO + DyCo-RLModel Scale=Qwen2.5-VL-3B2026.06 | — | 41 | — | |
| GRPO + DyCo-RLModel Scale=Qwen2.5-VL-7B2026.06 | — | 48.5 | — | |
| GSPOModel Scale=Qwen2.5-VL-3B2026.06 | — | 35.8 | — | |
| GSPOModel Scale=Qwen2.5-VL-7B2026.06 | — | 46.5 | — | |
| GSPO + DyCo-RLModel Scale=Qwen2.5-VL-3B2026.06 | — | 37.4 | — | |
| GSPO + DyCo-RLModel Scale=Qwen2.5-VL-7B2026.06 | — | 48.3 | — | |
| InternVL2.5-8B*Model Category=Open-Source Models2026.04 | — | 38.3 | 44.7 | |
| MUPO-Thinker-3BModel Category=Our Models2026.04 | — | 42.8 | 50.3 | |
| MUPO-Thinker-7BModel Category=Our Models2026.04 | — | 50.6 | 61.5 | |
| Qwen2.5-VL-7BModel Category=Open-Source Models2026.04 | — | 42.6 | 59.3 | |
| R1-OneVision-7BModel Category=Open-Source Models2026.04 | — | 45.6 | 52.5 | |
| SAPOModel Scale=Qwen2.5-VL-3B2026.06 | — | 39.4 | — | |
| SAPOModel Scale=Qwen2.5-VL-7B2026.06 | — | 42.3 | — | |
| SAPO + DyCo-RLModel Scale=Qwen2.5-VL-3B2026.06 | — | 44.5 | — | |
| SAPO + DyCo-RLModel Scale=Qwen2.5-VL-7B2026.06 | — | 47.7 | — | |
| Vision-R1-7BModel Category=Open-Source Models2026.04 | — | 49.7 | 53.8 | |
| VLAA-Thinker-7BModel Category=Open-Source Models2026.04 | — | 48.5 | 54.5 | |
| Zero-ShotModel Scale=Qwen2.5-VL-3B2026.06 | — | 34.5 | — | |
| Zero-ShotModel Scale=Qwen2.5-VL-7B2026.06 | — | 44.5 | — |