Multimodal Math Reasoning on MathVision
86AccuracyQwen3.5-27B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5-27BMode=REASONING, Architecture=Dense, # Total Params=27B, # Activated Params=27B2026.04 | 86 | |
| EXAONE 4.5 33BMode=REASONING, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 75.2 | |
| Qwen3-VL-235B-A22BMode=Thinking, Architecture=MoE, # Total Params=236B, # Activated Params=23B2026.04 | 74.6 | |
| GPT-5 miniMode=REASONING: HIGH, Architecture=-, # Total Params=-, # Activated Params=-2026.04 | 71.9 | |
| GPT-5 highInput Modality=Multimodal, LLM-as-a-Judge=GPT-4o2025.11 | 71.6 | |
| Qwen3-VL-32BMode=Thinking, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 70.2 | |
| Seed-1.5-thinkingAccess Type=Closed-Source2025.12 | 68.7 | |
| Gemini 2.5 ProInput Modality=Multimodal, LLM-as-a-Judge=GPT-4o2025.11 | 63.3 | |
| o12026.03 | 60.3 | |
| o1Model Category=Closed-Source Models2026.03 | 60.3 | |
| Claude Sonnet 4.5Input Modality=Multimodal, LLM-as-a-Judge=GPT-4o2025.11 | 58.7 | |
| InternVL3.5-8BModel Scale=8B, Optimization Method=N/A2026.06 | 56.8 | |
| Gemini-2.5-Pro-ThinkingAccess Type=Closed-Source2025.12 | 55.3 | |
| Claude-4-SonnetAccess Type=Closed-Source2025.12 | 54.6 | |
| Qwen2.5VL-72B x Qwen3-32BParam (B)=1042025.09 | 52.6 | |
| GPT-4.1Access Type=Closed-Source2025.12 | 51.8 | |
| Qwen3-VLModel Size=4B, Training Stage=instruct2026.05 | 51.6 | |
| InternVL3.5-4B (+GNDPO)Model Scale=4B, Optimization Method=GNDPO2026.06 | 51.6 | |
| Vision-R1Backbone=Vision-R1 (7B)2026.05 | 51.2 | |
| GLM-4v-Plus-20250111Param (B)=-2025.09 | 51.1 | |
| InternVL3.5-4B (+OPD)Model Scale=4B, Optimization Method=OPD2026.06 | 50.5 | |
| Qwen3-VL-SegModel Size=4B, Training Stage=S-22026.05 | 50.4 | |
| InternVL3.5-4B (+GSPO)Model Scale=4B, Optimization Method=GSPO2026.06 | 48.7 | |
| Gemini-2.0-ProModel Category=Proprietary Models2025.06 | 48.1 | |
| Qwen3-VLModel Size=4B, Training Stage=S-12026.05 | 47.9 | |
| Sora-2 AudioInput Modality=Audio, LLM-as-a-Judge=GPT-4o2025.11 | 46.7 | |
| Qwen2.5VL-7B x Qwen3-32BParam (B)=392025.09 | 46.6 | |
| GPT-4.1-20250414Param (B)=-2025.09 | 45.1 | |
| Sora-2 Last FrameInput Modality=Last Frame, LLM-as-a-Judge=GPT-4o2025.11 | 44.9 | |
| PDCRBackbone=Qwen2.5-VL-7B2026.05 | 44.8 | |
| Revisual-R1Access Type=Open-Source, Model Scale=7B2025.12 | 44.7 | |
| PACRBackbone=Qwen2.5-VL-7B2026.05 | 44.7 | |
| OursAccess Type=Open-Source, Model Scale=7B2025.12 | 44.2 | |
| DAPOBackbone=Qwen2.5-VL-7B2026.05 | 44.2 | |
| ChatGPT-4o-latestParam (B)=-2025.09 | 43.8 | |
| GPT-4oModel Category=Proprietary Models2025.06 | 43.8 | |
| Qwen2.5VL-7B x Qwen3-4BParam (B)=112025.09 | 43.2 | |
| Claude3.7-SonnetParam (B)=-2025.09 | 41.9 | |
| Claude-3.7-SonnetModel Category=Proprietary Models2025.06 | 41.9 | |
| Claude3.7-Sonnet2026.03 | 41.3 | |
| Gemini2-flash2026.03 | 41.3 | |
| Claude3.7-SonnetModel Category=Closed-Source Models2026.03 | 41.3 | |
| Gemini-2-flashModel Category=Closed-Source Models2026.03 | 41.3 | |
| Claude-3.7-SonnetModel Category=Close-source Models2026.04 | 41.3 | |
| GRPOBackbone=Qwen2.5-VL-7B2026.05 | 41.3 | |
| InternVL3.5-4B (Base (Instruct))Model Scale=4B, Optimization Method=Base (Instruct)2026.06 | 40.5 | |
| Zero-shot InferenceBackbone=Qwen2.5-VL-7B, Mode=Zero-shot2026.05 | 40.3 | |
| Qwen-2.5-VL-32B2026.03 | 40.1 | |
| Qwen-2.5-VL-32BModel Category=Open-Source General Models2026.03 | 40.1 | |
| PDCRBackbone=Qwen2.5-VL-3B2026.05 | 40.1 | |
| MM-Eureka-32B-R-TAPModel Category=Open-Source Reasoning Models2026.03 | 39.9 | |
| Gemma-3-27BParameters=27B2025.12 | 39.8 | |
| GRPOBackbone=Qwen2.5-VL-3B2026.05 | 39.4 | |
| Qwen2.5-VL-72BParam (B)=722025.09 | 39.3 | |
| Qwen2.5-VL-72BModel Category=Open-source Models (>70B)2025.06 | 39.3 | |
| DAPOBackbone=Qwen2.5-VL-3B2026.05 | 39.1 | |
| PACRBackbone=Qwen2.5-VL-3B2026.05 | 39 | |
| InternVL3-78BParam (B)=782025.09 | 38.8 | |
| Zero-shot InferenceBackbone=Qwen2.5-VL-3B, Mode=Zero-shot2026.05 | 38.6 | |
| InternVL3.5-2B (+GNDPO)Model Scale=2B, Optimization Method=GNDPO2026.06 | 38.4 | |
| Qwen2.5-VL-72BModel Category=Larger MLLMs without Reasoning2025.12 | 38.1 | |
| Qwen-2.5-VL-72B2026.03 | 38.1 | |
| Qwen-2.5-VL-72BModel Category=Open-Source General Models2026.03 | 38.1 | |
| InternVL3.5-2B (+OPD)Model Scale=2B, Optimization Method=OPD2026.06 | 37.9 | |
| GPT-4oModel Category=Larger MLLMs without Reasoning2025.12 | 36.5 | |
| Visionary-R1Backbone=Visionary-R1 (3B)2026.05 | 36.5 | |
| InternVL2.5-78B-MPOModel Category=Open-source Models (>70B)2025.06 | 36.2 | |
| RISEBackbone=Qwen3-VL-8B-Instruct, Self-evolving steps=602026.05 | 36.1 | |
| QVQ-72B-Preview2026.03 | 35.9 | |
| QVQ-72B-PreviewModel Category=Open-Source Reasoning Models2026.03 | 35.9 | |
| InternVL3.5-2B (+GSPO)Model Scale=2B, Optimization Method=GSPO2026.06 | 35.9 | |
| Perception-R1Backbone=Perception-R1 (7B)2026.05 | 35.7 | |
| Qwen2.5-VL-7B-Instruct + R-TAPModel Size=7B2026.03 | 35.3 | |
| QVQ-72B-PreviewParam (B)=722025.09 | 34.9 | |
| Qwen2.5-VL-7B + SRPOBackbone=Qwen2.5-VL-7B2026.05 | 34.54 | |
| MM-Eureka-32BModel Category=Open-Source Reasoning Models2026.03 | 34.4 | |
| ThinkLite-VLAccess Type=Open-Source, Model Scale=7B2025.12 | 32.9 | |
| RISEBackbone=Qwen3-VL-8B-Instruct, Self-evolving steps=402026.05 | 32.8 | |
| PeBR-R1Access Type=Open-Source, Model Scale=7B2025.12 | 32.7 | |
| AStar (Qwen2.5-7B)OS Only=✓, Training-Free=✓, Prior Data=0.5K, Pre. Time=50 mins2025.02 | 32.7 | |
| VL-RethinkerAccess Type=Open-Source, Model Scale=7B2025.12 | 32.3 | |
| InternVL2.5-38B-MPO2026.03 | 32.3 | |
| InternVL2.5-38B-MPOModel Category=Open-Source Reasoning Models2026.03 | 32.3 | |
| InternVL2.5-VL-78B2026.03 | 32.2 | |
| InternVL2.5-VL-78BModel Category=Open-Source General Models2026.03 | 32.2 | |
| InternVL2.5-78BModel Category=Open-source Models (>70B)2025.06 | 32.2 | |
| DIVA-GRPO-7B2026.03 | 32.1 | |
| VAPO-ThinkerAccess Type=Open-Source, Model Scale=7B2025.12 | 31.9 | |
| InternVL2.5-VL-38B2026.03 | 31.8 | |
| InternVL2.5-VL-38BModel Category=Open-Source General Models2026.03 | 31.8 | |
| MM-Eureka-7B-R-TAPModel Category=Open-Source Reasoning Models2026.03 | 31.7 | |
| ToR-DAPOData Size=39K (ViRL-39K), Base Model=Qwen-2.5-VL-7B2026.03 | 31.6 | |
| Vision-G1Access Type=Open-Source, Model Scale=7B2025.12 | 31.3 | |
| RISEBackbone=Qwen3-VL-8B-Instruct, Self-evolving steps=202026.05 | 31.28 | |
| GPT-4oParam (B)=-2025.09 | 31.2 | |
| PRCO-7BBackbone=Qwen2.5-VL-7B2026.05 | 30.92 | |
| Qwen2.5-VL-7B + GRPOBackbone=Qwen2.5-VL-7B2026.05 | 30.92 | |
| SelfJudgeTraining Data=Geo3K, Backbone=Qwen2.5-VL-7B2026.03 | 30.9 | |
| R1-OnevisionParam (B)=72025.09 | 30.6 | |
| VL-Rethinker-7BBackbone=Qwen2.5-VL-7B2026.05 | 30.59 |