Mathematical Reasoning on MathVista
229.2ScoreQwen-VL-7B-Chat
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-VL-7B-ChatModel Size=7B, Mode=Chat2023.11 | 229.2 | |
| EcoAlignCategory=Closed-source, Model=Qwen-VL-Max, Avg. Cost=12.72025.11 | 90.7 | |
| EcoAlignCategory=Closed-source, Model=Gemini-2.5-Flash, Avg. Cost=24.12025.11 | 89.6 | |
| CoTCategory=Closed-source, Model=Qwen-VL-Max, Avg. Cost=1142025.11 | 89.5 | |
| CoDCategory=Closed-source, Model=Qwen-VL-Max, Avg. Cost=34.32025.11 | 89.1 | |
| CoDCategory=Closed-source, Model=Gemini-2.5-Flash, Avg. Cost=27.52025.11 | 88.5 | |
| CoTCategory=Closed-source, Model=Gemini-2.5-Flash, Avg. Cost=86.62025.11 | 88.2 | |
| CoTCategory=Closed-source, Model=GPT-4o, Avg. Cost=104.32025.11 | 86.8 | |
| EcoAlignCategory=Open-source, Model=InternVL3-14B, Avg. Cost=39.32025.11 | 86 | |
| EcoAlignCategory=Closed-source, Model=GPT-4o, Avg. Cost=21.22025.11 | 85.4 | |
| CoDCategory=Closed-source, Model=GPT-4o, Avg. Cost=25.82025.11 | 84.2 | |
| CoDCategory=Open-source, Model=InternVL3-14B, Avg. Cost=45.52025.11 | 84.1 | |
| CoTCategory=Open-source, Model=InternVL3-14B, Avg. Cost=134.22025.11 | 83.9 | |
| LongCat-Next2026.03 | 83.1 | |
| Gemini-2.5-ProDate=2024.12, Tokenization Paradigm=Proprietary Models2026.02 | 82.7 | |
| Gemini-2.5-ProModel Category=Proprietary Models2026.06 | 82.7 | |
| GPT-5Inference mode=Thinking mode2026.04 | 81.9 | |
| MiMo-VL-RLLLM Backbone=MiMo-7B, Open Data=false2025.12 | 81.5 | |
| Qwen3-VLInference mode=Thinking mode, Size=8B2026.04 | 81.4 | |
| GPT-5-highModel Category=Proprietary Models2026.06 | 81.3 | |
| MiniCPM-o 4.5Inference mode=Thinking mode, Size=9B2026.04 | 81 | |
| VL-Scaler-MiMOSize=7B2026.05 | 80.3 | |
| MiniCPM-o 4.5Size=9B, mode=instruct2026.04 | 80.1 | |
| Qwen3-OmniInference mode=Thinking mode, Size=30B-A3B2026.04 | 80 | |
| MetaForgeTool Setting=w/ OOD Tools2026.06 | 79.78 | |
| Gemini 2.5 FlashInference mode=Thinking mode2026.04 | 79.4 | |
| BaseCategory=Closed-source, Model=Qwen-VL-Max, Avg. Cost=12025.11 | 79 | |
| MiMO-VL-InstructSize=7B2026.05 | 78.8 | |
| InternVL3.5Size=8B, mode=instruct2026.04 | 78.4 | |
| OPDModel Scale=8B2026.06 | 78.1 | |
| MetaForgeTool Setting=w/ IID Tools2026.06 | 78.01 | |
| BaseCategory=Closed-source, Model=Gemini-2.5-Flash, Avg. Cost=12025.11 | 77.7 | |
| Qwen3-VLSize=8B, mode=instruct2026.04 | 77.2 | |
| InternVL3.5Model Scale=4B, Date=2025.08, Tokenization Paradigm=Understanding-only Models2026.02 | 77.1 | |
| OPD+ViCuRModel Scale=8B2026.06 | 76.7 | |
| GRPOModel Scale=8B2026.06 | 76.6 | |
| TVI-CoTModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=8B2026.06 | 76.6 | |
| KelixModel Scale=8B, Tokenization Paradigm=Discrete Tokenization2026.02 | 76.5 | |
| LVRPOType=Unified (>1.5B), # LLM Params=7B2026.03 | 76.2 | |
| VLM-GuardCategory=Closed-source, Model=Qwen-VL-Max, Avg. Cost=2.22025.11 | 76.1 | |
| Qwen3-OmniSize=30B-A3B, mode=instruct2026.04 | 75.9 | |
| Base ModelModel Scale=8B2026.06 | 75.8 | |
| VAPO-Thinker-7BModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=7B2026.06 | 75.6 | |
| Gemini 2.5 FlashSize=-, mode=instruct2026.04 | 75.3 | |
| BiPS-General-7BData=13K+39K2025.12 | 75 | |
| VL-RethinkerSize=7B2026.05 | 74.9 | |
| VL-Rethinker-7BModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=7B2026.06 | 74.9 | |
| Qwen2.5-VL-InstructSize=72B2026.05 | 74.8 | |
| Qwen3-VL-4B (Base)Base Model=Qwen3-VL-4B, Alignment Strategy=Base2026.06 | 74.8 | |
| VLsIParameters=7B2024.12 | 74.7 | |
| VLSI-7BSize=7B2024.12 | 74.7 | |
| Claude-3.7-SonnetData=-2025.12 | 74.5 | |
| Claude-Opus-4.1Model Category=Proprietary Models2026.06 | 74.5 | |
| GRPOData=13K + 39K2025.12 | 74.3 | |
| OPSD+ViCuRModel Scale=8B2026.06 | 74.1 | |
| VL-ScalerSize=7B2026.05 | 73.8 | |
| BiPS-Chart-7BData=13K2025.12 | 73.5 | |
| BaseCategory=Open-source, Model=InternVL3-14B, Avg. Cost=12025.11 | 73.5 | |
| OPSDModel Scale=8B2026.06 | 73.5 | |
| Vision-R1-7BModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=7B2026.06 | 73.5 | |
| ManzanoModel Scale=30B, Date=2025.09, Tokenization Paradigm=Hybrid Tokenization2026.02 | 73.3 | |
| Vision-R1-7BData=73K + 137K2025.12 | 73.2 | |
| Qwen3-VL-8B (Baseline)Model Category=Baseline, Model Scale=8B2026.06 | 73.2 | |
| BagelModel Scale=14B, Date=2025.05, Tokenization Paradigm=Continuous Tokenization2026.02 | 73.1 | |
| Bagel#LLM=14B2025.12 | 73.1 | |
| BaseCategory=Closed-source, Model=GPT-4o, Avg. Cost=12025.11 | 73.1 | |
| VLM-GuardCategory=Closed-source, Model=Gemini-2.5-Flash, Avg. Cost=1.42025.11 | 73.1 | |
| BAGEL2026.03 | 73.1 | |
| BAGELType=Unified (>1.5B), # LLM Params=7B2026.03 | 73.1 | |
| BAGALType=Und., # Params=14B2025.03 | 73.1 | |
| COPSD (Standard)Base Model=Qwen3-VL-4B, Alignment Strategy=COPSD (Standard)2026.06 | 72.7 | |
| COPSD (Hybrid)Base Model=Qwen3-VL-4B, Alignment Strategy=COPSD (Hybrid)2026.06 | 72.5 | |
| PRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 72.2 | |
| Qwen3-VLScale=4B2026.05 | 72 | |
| MM-Eureka-7BModel Scale=7B, Rollout=82026.06 | 71.9 | |
| InternVL3.5Number of Parameters=2.3B2025.12 | 71.8 | |
| CodeV-7B-RLTool use=Yes, Training=RL2025.11 | 71.8 | |
| InternVL3LLM Backbone=Qwen2.5-7B, Open Data=false2025.12 | 71.6 | |
| InternVL3-8BModel Category=Open-Source MLLMs, Model Scale=8B2026.06 | 71.6 | |
| InternVL3Model Size=8B2026.06 | 71.6 | |
| Qwen3.5-VLScale=2B2026.05 | 71.4 | |
| ThinkLite-7BTool use=No2025.11 | 71.3 | |
| Pixel-Reasoner-7BTool use=Yes2025.11 | 71.2 | |
| SAIL-VL2Number of Parameters=2.7B2025.12 | 71.1 | |
| Phantom-7BSize=7B2024.12 | 70.9 | |
| PhantomParameters=7B2024.09 | 70.9 | |
| DeepEyes-7BData=14K + 33K2025.12 | 70.8 | |
| VL-Rethinker-7BModel Scale=7B, Rollout=82026.06 | 70.6 | |
| VLM-GuardCategory=Open-source, Model=InternVL3-14B, Avg. Cost=1.92025.11 | 70.5 | |
| VPPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 70.5 | |
| InternVL3-8BData=-2025.12 | 70.4 | |
| InternVL3-8BTool use=No2025.11 | 70.4 | |
| Latent DenoisingArchitecture=Qwen-2.5-VL2026.04 | 70.3 | |
| OpenVLThinkerSize=7B2026.05 | 70.2 | |
| Qwen2-VL-72BSize=72B2024.12 | 69.7 | |
| VLM-GuardCategory=Closed-source, Model=GPT-4o, Avg. Cost=3.12025.11 | 69.7 | |
| LLaVA-OneVision-1.5-8BModel Category=Open-Source MLLMs, Model Scale=8B2026.06 | 69.6 | |
| Ovis-U1#LLM=1.5B2025.12 | 69.4 | |
| Ovis-U12026.03 | 69.4 | |
| SFTBase Model=Qwen3-VL-4B, Alignment Strategy=SFT2026.06 | 69.4 |