Mathematical Reasoning on MathVista mini
86.2AccuracyQwen3.5
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5Parameters=35BA3B, Internal Reasoning (Think mode)=false2026.05 | 86.2 | |
| Qwen3.5Parameters=9B, Internal Reasoning (Think mode)=false2026.05 | 85.7 | |
| SenseNova-U1Parameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 85.3 | |
| SenseNova-U1Parameters=8B, Internal Reasoning (Think mode)=true2026.05 | 84.2 | |
| Qwen2-VL-72B-ThinkingMax Output Tokens=4K2026.03 | 83.9 | |
| RLR³Training Source=OpenMMR2026.05 | 83.9 | |
| Qwen2-VL-72B-ThinkingMax Output Tokens=40K2026.03 | 83.8 | |
| RLR³Training Source=ViRL2026.05 | 83.7 | |
| RLVRTraining Source=OpenMMR2026.05 | 83.5 | |
| LongCat-NextParameters=68BA3B, Internal Reasoning (Think mode)=false2026.05 | 83.1 | |
| RLVRTraining Source=DeepVision2026.05 | 83 | |
| RLR³Training Source=DeepVision2026.05 | 82.8 | |
| RLVRTraining Source=ViRL2026.05 | 82.3 | |
| Qwen3VLParameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 81.9 | |
| Official thinkingSource=Qwen3-VL2026.05 | 81.9 | |
| Qwen-8B-DeltaThinker2026.05 | 81.74 | |
| Qwen3VLParameters=8B, Internal Reasoning (Think mode)=true2026.05 | 81.4 | |
| Qwen3-VL-8B-Thinking2026.05 | 81.12 | |
| OPD w/ DELTAPROMPTS (10k)Training Data=DELTAPROMPTS (10k)2026.05 | 81.05 | |
| Individual Geometry2025.05 | 80.34 | |
| Qwen3-VL-30B A3B-InstructParameters=30B A3B, Type=Instruct2025.12 | 80.1 | |
| Official instructSource=Qwen3-VL2026.05 | 80.1 | |
| OptMerge2025.05 | 80.01 | |
| GLM-4.6V FlashParameters=4.6B, Variant=Flash2025.12 | 80 | |
| Base instruct2026.05 | 79.8 | |
| Qwen2-VL-7B-ThinkingMax Output Tokens=40K2026.03 | 79.5 | |
| Qwen2.5-VL-32B-Instruct2025.05 | 79.21 | |
| GPT-5 miniVersion=high2026.05 | 79.1 | |
| OPD w/ Seed Data (10k)Training Data=Seed Data (10k)2026.05 | 79.05 | |
| Individual Grounding2025.05 | 78.86 | |
| Kimi-VL-A3B-Thinking2026.03 | 78.6 | |
| Individual VQA2025.05 | 78.4 | |
| Qwen2-VL-7B-ThinkingMax Output Tokens=4K2026.03 | 77.7 | |
| Qwen3-VL-8BFine-tuning setting=CPO2026.07 | 76 | |
| Innovator-VLvariant=8B-Thinking, parameters=8B2026.01 | 75.3 | |
| Individual OCR2025.05 | 75.28 | |
| Phi-4-reasoning-vision-15BThinking Mode=default2026.03 | 75.2 | |
| Qwen2.5-VLNumber of Parameters=72B2025.02 | 74.8 | |
| Qwen3-VL 8B-InstructParameters=8B, Type=Instruct2025.12 | 74.6 | |
| LUSPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=LUSPO2026.02 | 74.4 | |
| Step-GUI-8BParameters=8B2025.12 | 74.4 | |
| Individual Chart2025.05 | 74.38 | |
| Innovator-VLvariant=8B-Instruct, parameters=8B2026.01 | 74.3 | |
| Phi-4-reasoning-vision-15BThinking Mode=force thinking2026.03 | 74.1 | |
| GSPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=GSPO2026.02 | 73.9 | |
| Qwen3-VLparameters=8B2026.01 | 73.3 | |
| GRPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=GRPO2026.02 | 72.8 | |
| Gemma4Parameters=26BA4B, Internal Reasoning (Think mode)=false2026.05 | 72.7 | |
| InternVL3.5parameters=8B2026.01 | 72.4 | |
| InternVL2.5Number of Parameters=78B2025.02 | 72.3 | |
| Step-GUI-4BParameters=4B2025.12 | 72.3 | |
| InternVL2.5-78Bparameters=78B2024.12 | 72.3 | |
| InternVL2.5-38Bparameters=38B2024.12 | 71.9 | |
| Qwen2-VLNumber of Parameters=72B2025.02 | 70.5 | |
| Qwen2-VL-72Bparameters=72B2024.12 | 70.5 | |
| Gemini-2.5 flash-liteVariant=flash-lite2025.12 | 70.3 | |
| OursModel=Kimi-VL A3B-Thinking2025.10 | 69.78 | |
| Intern-S1variant=mini, parameters=9B2026.01 | 68.4 | |
| Qwen2.5-VLNumber of Parameters=7B2025.02 | 68.2 | |
| Claude-3.5 Sonnet-0620Model Version=06202025.02 | 67.7 | |
| InternVL2.5-26Bparameters=26B2024.12 | 67.7 | |
| Claude-3.5-Sonnet2024.12 | 67.7 | |
| LLaVA-OneVision-72Bparameters=72B2024.12 | 67.5 | |
| Qwen2.5-VL-7B-Instruct (w/o RLVR)Base model=Qwen2.5-VL-7B-Instruct, Algorithm=w/o RLVR2026.02 | 67.4 | |
| AGLAModel=Kimi-VL A3B-Thinking2025.10 | 67.32 | |
| Ovis1.6-Gemma2-9Bparameters=9B, backbone=Gemma22024.12 | 67.2 | |
| NVLM-D-72Bparameters=72B2024.12 | 66.6 | |
| MiMo-VLtraining=7B-RL, parameters=7B2026.01 | 65.9 | |
| InternVL2-Llama3-76Bparameters=76B, backbone=Llama32024.12 | 65.5 | |
| MiMo-VLtraining=7B-SFT, parameters=7B2026.01 | 65.4 | |
| InfiniteVL-4BComplexity=O(1)2025.12 | 65.4 | |
| Qwen3-VL-8BFine-tuning setting=GSPO2026.07 | 64.9 | |
| InternVL2.5-8Bparameters=8B2024.12 | 64.4 | |
| Qwen3-VL-8BFine-tuning setting=LoRA2026.07 | 64.4 | |
| Gemini-1.5-Pro2024.12 | 63.9 | |
| GPT-4o 0513Model Version=05132025.02 | 63.8 | |
| GPT-4o-202405132024.12 | 63.8 | |
| InternVL2-40Bparameters=40B2024.12 | 63.7 | |
| VCDModel=Kimi-VL A3B-Thinking2025.10 | 63.51 | |
| VanillaModel=Kimi-VL A3B-Thinking2025.10 | 63.48 | |
| LLaVA-OVversion=1.5, parameters=8B2026.01 | 63.3 | |
| Qwen2.5-VLNumber of Parameters=3B2025.02 | 62.3 | |
| Qwen2.5VL-3BComplexity=O(n)2025.12 | 62.3 | |
| MiniCPM-Vversion=4.5, parameters=8B2026.01 | 61.1 | |
| Qwen3-VL-8BFine-tuning setting=GRPO2026.07 | 61.1 | |
| MiniCPM-V2.62024.12 | 60.6 | |
| InternVL2.5-4BComplexity=O(n)2025.12 | 60.5 | |
| InternVL2.5-4Bparameters=4B2024.12 | 60.5 | |
| AGLAModel=R1-Onevision 7B2025.10 | 60.21 | |
| OursModel=R1-Onevision 7B2025.10 | 60.09 | |
| VanillaModel=R1-Onevision 7B2025.10 | 59.92 | |
| VCDModel=R1-Onevision 7B2025.10 | 59.69 | |
| GPT-5 miniVersion=minimal2026.05 | 59.6 | |
| InternVL2-26Bparameters=26B2024.12 | 59.4 | |
| OursModel=Ocean-R1 7B Instruct2025.10 | 59.32 | |
| Aquila-VL-2Bparameters=2B2024.12 | 59 | |
| InternVL2-4Bparameters=4B2024.12 | 58.6 | |
| Molmo-72Bparameters=72B2024.12 | 58.6 | |
| InternVL2-8Bparameters=8B2024.12 | 58.3 | |
| Qwen2-VL-7Bparameters=7B2024.12 | 58.2 |