Mathematical Reasoning on MathVerse mini
82.6AccuracyQwen3-VL 32B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL 32BModel Source Category=Open-weight VLMs, Model Parameters=32B, Evaluation Mode=Thinking Mode2026.01 | 82.6 | |
| MMFineReason-8BModel Source Category=Ours, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 81.5 | |
| Qwen3-VL 30B-A3BModel Source Category=Open-weight VLMs, Model Parameters=30B-A3B, Evaluation Mode=Thinking Mode2026.01 | 79.6 | |
| RLR³Training Source=DeepVision2026.05 | 79.6 | |
| Official thinkingSource=Qwen3-VL2026.05 | 79.6 | |
| GPT5 miniModel Source Category=Closed-source VLMs, Evaluation Mode=Thinking Mode2026.01 | 78.8 | |
| GPT-5 miniVersion=high2026.05 | 78.8 | |
| MMFineReason-4BModel Source Category=Ours, Model Parameters=4B, Evaluation Mode=Thinking Mode2026.01 | 78.7 | |
| Qwen2-VL-72B-ThinkingMax Output Tokens=4K2026.03 | 78.3 | |
| Qwen2-VL-72B-ThinkingMax Output Tokens=40K2026.03 | 78.2 | |
| Gemini-2.5 FlashModel Source Category=Closed-source VLMs, Evaluation Mode=Thinking Mode2026.01 | 77.7 | |
| Qwen3-VL 8BModel Source Category=Open-weight VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 77.7 | |
| RLR³Training Source=OpenMMR2026.05 | 76.9 | |
| RLR³Training Source=ViRL2026.05 | 76.6 | |
| RLVRTraining Source=OpenMMR2026.05 | 75.5 | |
| RLVRTraining Source=DeepVision2026.05 | 73.9 | |
| Qwen2-VL-7B-ThinkingMax Output Tokens=40K2026.03 | 73.3 | |
| RLVRTraining Source=ViRL2026.05 | 72.1 | |
| Base instruct2026.05 | 71.7 | |
| Official instructSource=Qwen3-VL2026.05 | 70.2 | |
| MMFineReason-2BModel Source Category=Ours, Model Parameters=2B, Evaluation Mode=Thinking Mode2026.01 | 69.2 | |
| MMR1 8BModel Source Category=Open-source VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 67.3 | |
| Qwen2-VL-7B-ThinkingMax Output Tokens=4K2026.03 | 67.3 | |
| OMR 7BModel Source Category=Open-source VLMs, Model Parameters=7B, Evaluation Mode=Thinking Mode2026.01 | 63.8 | |
| Kimi-VL-A3B-Thinking2026.03 | 61 | |
| HoneyBee 8BModel Source Category=Open-source VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 60.9 | |
| Qwen3-VL-8B-InstructParameters=8B, Instruct=true2026.05 | 60.9 | |
| Innovator-VLvariant=8B-Thinking, parameters=8B2026.01 | 58.73 | |
| Qwen2.5-VLNumber of Parameters=72B2025.02 | 57.6 | |
| Qwen3-VL-8BFine-tuning setting=CPO2026.07 | 57.51 | |
| Qwen3-VLparameters=8B2026.01 | 57.28 | |
| CAVE-7BParameters=7B2026.05 | 55.7 | |
| Innovator-VLvariant=8B-Instruct, parameters=8B2026.01 | 54.77 | |
| MiMo-VLtraining=7B-SFT, parameters=7B2026.01 | 54.14 | |
| InternVL3.5-8BParameters=8B2026.05 | 53.2 | |
| Phi-4-reasoning-vision-15BThinking Mode=force thinking2026.03 | 53.1 | |
| InternVL2.5 (Previous SoTA)reference=Chen et al. (2024d)2025.02 | 51.7 | |
| GPT-4o 0513Model Version=05132025.02 | 51.7 | |
| InternVL2.5-78Bparameters=78B2024.12 | 51.7 | |
| MiMo-VLtraining=7B-RL, parameters=7B2026.01 | 50.28 | |
| Claude-3.5 Sonnet-0620Model Version=06202025.02 | 50.2 | |
| GPT-4o-202405132024.12 | 50.2 | |
| GPT-4oTool use=No2025.11 | 50.2 | |
| InternVL2.5-38Bparameters=38B2024.12 | 49.4 | |
| Qwen2.5-VLNumber of Parameters=7B2025.02 | 49.2 | |
| CodeV-7B-RLTool use=Yes, Training=RL2025.11 | 49.2 | |
| DeepEyes-7BParameters=7B2026.05 | 49.2 | |
| InternVL3.5parameters=8B2026.01 | 48.38 | |
| ThinkLite-7BTool use=No2025.11 | 48.2 | |
| Qwen2.5-VLParameters=7B2026.05 | 48.1 | |
| Qwen2.5-VLNumber of Parameters=3B2025.02 | 47.6 | |
| Pixel-Reasoner-7BTool use=Yes2025.11 | 46.9 | |
| Qwen2.5-VL-7BTool use=No2025.11 | 45.5 | |
| DeepEyes-7BTool use=Yes2025.11 | 45.4 | |
| Phi-4-reasoning-vision-15BThinking Mode=default2026.03 | 44.9 | |
| CodeV-7B-SFTTool use=Yes, Training=SFT2025.11 | 44.2 | |
| Qwen3-VL-8BFine-tuning setting=GRPO2026.07 | 44.19 | |
| LLaVA-OVversion=1.5, parameters=8B2026.01 | 43.88 | |
| Thyme-RL-7BTool use=Yes2025.11 | 43.6 | |
| Qwen3-VL-8BFine-tuning setting=Base2026.07 | 43.17 | |
| MiniCPM-Vversion=4.5, parameters=8B2026.01 | 42.92 | |
| Qwen3-VL-8BFine-tuning setting=GSPO2026.07 | 42.84 | |
| InternVL2-Llama3-76Bparameters=76B, backbone=Llama32024.12 | 42.8 | |
| InternVL2.5-26Bparameters=26B2024.12 | 40.1 | |
| InternVL3-8BTool use=No2025.11 | 40.1 | |
| Intern-S1variant=mini, parameters=9B2026.01 | 39.97 | |
| InternVL2.5-8Bparameters=8B2024.12 | 39.5 | |
| LLaVA-OneVision-72Bparameters=72B2024.12 | 39.1 | |
| InternVL2.5-4Bparameters=4B2024.12 | 37.1 | |
| InternVL2-8Bparameters=8B2024.12 | 37 | |
| GPT-5 miniVersion=minimal2026.05 | 36.5 | |
| InternVL2-40Bparameters=40B2024.12 | 36.3 | |
| Qwen3-VL-8BFine-tuning setting=LoRA2026.07 | 34.92 | |
| GPT-4V2024.12 | 32.8 | |
| InternVL2-4Bparameters=4B2024.12 | 32 | |
| Qwen2-VL-7Bparameters=7B2024.12 | 31.9 | |
| InternVL2-26Bparameters=26B2024.12 | 31.1 | |
| InternVL2.5-2Bparameters=2B2024.12 | 30.6 | |
| gemma-2-9b-it2026.03 | 29.8 | |
| InternVL-Chat-V1.52024.12 | 28.4 | |
| Qwen3-VL-8BFine-tuning setting=FFT2026.07 | 28.22 | |
| InternVL2.5-1Bparameters=1B2024.12 | 28 | |
| Aquila-VL-2Bparameters=2B2024.12 | 26.2 | |
| MiniCPM-V2.62024.12 | 25.7 | |
| InternVL2-2Bparameters=2B2024.12 | 25.3 | |
| Phi-3.5-Vision-4Bparameters=4B2024.12 | 24.1 | |
| Qwen2-VL-2Bparameters=2B2024.12 | 21 | |
| InternVL2-1Bparameters=1B2024.12 | 18.4 | |
| LLaVA-OneVision-0.5Bparameters=0.5B2024.12 | 17.9 |