Mathematical Reasoning on MathVision (Accuracy)
83.9AccuracyQwen3.5
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5Parameters=35BA3B, Internal Reasoning (Think mode)=false2026.05 | 83.9 | |
| Gemini3-proModel Category=Closed-Source SOTA Models2026.05 | 83.4 | |
| SenseNova-U1Parameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 79.63 | |
| Qwen3.5Parameters=9B, Internal Reasoning (Think mode)=false2026.05 | 78.9 | |
| SenseNova-U1Parameters=8B, Internal Reasoning (Think mode)=true2026.05 | 75.82 | |
| Qwen3-VL-235B-ThinkingModel Category=Open-Source Large Baselines2026.05 | 74.6 | |
| GPT-5-20250807Model Category=Closed-Source SOTA Models2026.05 | 72 | |
| Gemma4Parameters=26BA4B, Internal Reasoning (Think mode)=false2026.05 | 68.95 | |
| Doubao-Seed-1.6Model Category=Closed-Source SOTA Models2026.05 | 67.8 | |
| Qwen3VLParameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 65.7 | |
| GLM-4.5VModel Category=Closed-Source SOTA Models2026.05 | 65.6 | |
| LongCat-NextParameters=68BA3B, Internal Reasoning (Think mode)=false2026.05 | 64.7 | |
| Qwen3VLParameters=8B, Internal Reasoning (Think mode)=true2026.05 | 62.7 | |
| OSTBase Model=Qwen3-VL-8B-Instruct, Sampling Ratio=Best-20%2026.05 | 58.5 | |
| DEITABase Model=Qwen3-VL-8B-Instruct, Sampling Ratio=Top-20%2026.05 | 58.3 | |
| Qwen3-VL-8B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 58.2 | |
| Qwen3-VL-8B-InstructStrategy=Base2026.05 | 57.8 | |
| Qwen3-VL-8B-Instruct + RandomSampling Ratio=20%2026.05 | 57.1 | |
| Kimi-vl-A3B-thinkingModel Category=Open-Source Large Baselines2026.05 | 56.8 | |
| OSTBase Model=Qwen3-VL-4B-Instruct, Sampling Ratio=Best-20%2026.05 | 54.8 | |
| Qwen3-VL-8B-Instruct + Full SFTSampling Ratio=100%2026.05 | 54.8 | |
| Qwen3-VL-4B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 54.6 | |
| DEITABase Model=Qwen3-VL-4B-Instruct, Sampling Ratio=Top-20%2026.05 | 54.5 | |
| Qwen3-VL-4B-InstructStrategy=Base2026.05 | 54.1 | |
| Qwen3-VL-4B-Instruct + RandomSampling Ratio=20%2026.05 | 54 | |
| Qwen3-VL-4B-Instruct + Full SFTSampling Ratio=100%2026.05 | 53.8 | |
| Qwen3-VL-8BFine-tuning setting=CPO2026.07 | 51.61 | |
| Gemini-2.0-FlashActivation Replay=false2025.11 | 47.8 | |
| Qwen3-VL-4B-Instruct / GRPOStudent=Qwen3-VL-4B-Instruct, Teacher=Self, Method=GRPO2026.07 | 46.7 | |
| Qwen3-VL-4B-Instruct / H-OPD (8B)Student=Qwen3-VL-4B-Instruct, Teacher=Qwen3-VL-8B + Qwen3-8B, Method=H-OPD2026.07 | 46.7 | |
| Qwen3-VL-4B-Instruct / ExOPD (8B)Student=Qwen3-VL-4B-Instruct, Teacher=Qwen3-VL-8B, Method=ExOPD2026.07 | 46.4 | |
| Qwen3-VL-4B-Instruct / OPD (8B)Student=Qwen3-VL-4B-Instruct, Teacher=Qwen3-VL-8B, Method=OPD2026.07 | 46.1 | |
| Qwen3-VL-8BType=Baseline, Evaluation Tool=VLMEvalKit2026.07 | 45.4 | |
| Qwen3-VL-4BType=Baseline, Evaluation Tool=VLMEvalKit2026.07 | 44.7 | |
| Qwen3-VL-4B-Instruct / SFT (8B)Student=Qwen3-VL-4B-Instruct, Teacher=Qwen3-VL-8B, Method=SFT2026.07 | 42.1 | |
| Claude-3.5 SonnetActivation Replay=false2025.11 | 41.3 | |
| Qwen3-VL-2B-Instruct / H-OPD (8B)Student=Qwen3-VL-2B-Instruct, Teacher=Qwen3-VL-8B + Qwen3-8B, Method=H-OPD2026.07 | 36.8 | |
| Qwen3-VL-8BFine-tuning setting=Base2026.07 | 36.78 | |
| MM-Eureka-Qwen-32BActivation Replay=true2025.11 | 35.5 | |
| MM-Eureka-Qwen-32BActivation Replay=false2025.11 | 35.2 | |
| Qwen3-VL-2B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 35 | |
| QvQ-72B-PreviewActivation Replay=false2025.11 | 34.9 | |
| DEITABase Model=Qwen3-VL-2B-Instruct, Sampling Ratio=Top-20%2026.05 | 34.9 | |
| OSTBase Model=Qwen3-VL-2B-Instruct, Sampling Ratio=Best-20%2026.05 | 34.8 | |
| TemplateRLBackbone=Qwen2.5-VL-3B-Instruct, Batch size=128, Samples per question=16, Template guidance quantity=22025.05 | 34.6 | |
| Qwen3-VL-2B-InstructStrategy=Base2026.05 | 34.5 | |
| Qwen3-VL-2B-Instruct + RandomSampling Ratio=20%2026.05 | 34.3 | |
| Qwen3-VL-2B-Instruct + Full SFTSampling Ratio=100%2026.05 | 34.1 | |
| MMR1-Math-v0-7BActivation Replay=true2025.11 | 33.9 | |
| Qwen3-VL-2B-Instruct / H-OPD (4B)Student=Qwen3-VL-2B-Instruct, Teacher=Qwen3-VL-4B + Qwen3-4B, Method=H-OPD2026.07 | 33.5 | |
| VL-Rethinker-7BActivation Replay=true2025.11 | 33.2 | |
| MMR1-Math-v0-7BActivation Replay=false2025.11 | 32.6 | |
| Qwen3-VL-8BFine-tuning setting=GRPO2026.07 | 31.78 | |
| Qwen3-VL-2B-Instruct / ExOPD (4B)Student=Qwen3-VL-2B-Instruct, Teacher=Qwen3-VL-4B, Method=ExOPD2026.07 | 31.6 | |
| MM-Eureka-Qwen-7BActivation Replay=true2025.11 | 31.5 | |
| GPT-4oActivation Replay=false2025.11 | 31.1 | |
| GRPOBackbone=Qwen2.5-VL-3B-Instruct, Batch size=128, Samples per question=162025.05 | 30.9 | |
| Qwen3-VL-2B-Instruct / OPD (4B)Student=Qwen3-VL-2B-Instruct, Teacher=Qwen3-VL-4B, Method=OPD2026.07 | 30.9 | |
| MM-Eureka-Qwen-7BActivation Replay=false2025.11 | 30.6 | |
| Qwen3-VL-2B-Instruct / ExOPD (8B)Student=Qwen3-VL-2B-Instruct, Teacher=Qwen3-VL-8B, Method=ExOPD2026.07 | 30.6 | |
| VL-Rethinker-7BActivation Replay=false2025.11 | 30.3 | |
| Qwen3-VL-2B-Instruct / GRPOStudent=Qwen3-VL-2B-Instruct, Teacher=Self, Method=GRPO2026.07 | 30.3 | |
| Qwen3-VL-2B-Instruct / OPD (8B)Student=Qwen3-VL-2B-Instruct, Teacher=Qwen3-VL-8B, Method=OPD2026.07 | 30.3 | |
| InternVL3-8BActivation Replay=false2025.11 | 28.6 | |
| Qwen3-VL-2BType=Baseline, Evaluation Tool=VLMEvalKit2026.07 | 28.6 | |
| Qwen3-VL-2B-Instruct / SFT (8B)Student=Qwen3-VL-2B-Instruct, Teacher=Qwen3-VL-8B, Method=SFT2026.07 | 28.2 | |
| CoTBackbone=Qwen2.5-VL-3B-Instruct, Batch size=128, Samples per question=162025.05 | 27.7 | |
| GPT-4o-miniActivation Replay=false2025.11 | 27.3 | |
| Qwen3-VL-2B-Instruct / SFT (4B)Student=Qwen3-VL-2B-Instruct, Teacher=Qwen3-VL-4B, Method=SFT2026.07 | 26.8 | |
| Qwen3-VL-8BFine-tuning setting=GSPO2026.07 | 26.35 | |
| +BETAPRMSelector=+BETAPRM, Backbone=InternVL3-14B, Best-of-N strategy=Best-of-162026.05 | 25.66 | |
| +BETAPRMSelector=+BETAPRM, Backbone=InternVL2.5-8B, Best-of-N strategy=Best-of-162026.05 | 25.66 | |
| Qwen2.5-VL-7BActivation Replay=false2025.11 | 25.5 | |
| +BETAPRMSelector=+BETAPRM, Backbone=InternVL3-8B, Best-of-N strategy=Best-of-162026.05 | 24.34 | |
| +BETAPRMSelector=+BETAPRM, Backbone=Qwen2.5-VL-7B, Best-of-N strategy=Best-of-162026.05 | 24.34 | |
| Qwen3-VL-8BFine-tuning setting=LoRA2026.07 | 23.98 | |
| +Standard PRMSelector=+Standard PRM, Backbone=InternVL3-14B, Best-of-N strategy=Best-of-162026.05 | 23.03 | |
| +Standard PRMSelector=+Standard PRM, Backbone=InternVL3-8B, Best-of-N strategy=Best-of-162026.05 | 22.69 | |
| Kimi-VL-16BActivation Replay=false2025.11 | 21.8 | |
| +Standard PRMSelector=+Standard PRM, Backbone=InternVL2.5-8B, Best-of-N strategy=Best-of-162026.05 | 21.38 | |
| +Standard PRMSelector=+Standard PRM, Backbone=Qwen2.5-VL-7B, Best-of-N strategy=Best-of-162026.05 | 21.38 | |
| +Base (w/o training)Selector=+Base (w/o training), Backbone=InternVL2.5-8B, Best-of-N strategy=Best-of-162026.05 | 20.72 | |
| +Base (w/o training)Selector=+Base (w/o training), Backbone=InternVL3-14B, Best-of-N strategy=Best-of-162026.05 | 19.74 | |
| Qwen3-VL-8BFine-tuning setting=FFT2026.07 | 18.82 | |
| +Base (w/o training)Selector=+Base (w/o training), Backbone=InternVL3-8B, Best-of-N strategy=Best-of-162026.05 | 18.75 | |
| Single PassSelector=Single Pass, Backbone=InternVL2.5-8B, Best-of-N strategy=1 (Single Pass)2026.05 | 18.08 | |
| LLaVA-OV-7BActivation Replay=false2025.11 | 17 | |
| Qwen2.5-VL-3BActivation Replay=false2025.11 | 16.1 | |
| +Base (w/o training)Selector=+Base (w/o training), Backbone=Qwen2.5-VL-7B, Best-of-N strategy=Best-of-162026.05 | 15.46 |