Mathematical Reasoning on MathVista (Accuracy)
89.2AccuracyGemini 3-Pro
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini 3-Pro2026.02 | 89.2 | — | — | |
| Gemini3-proModel Category=Closed-Source SOTA Models2026.05 | 88.5 | — | — | |
| Qwen3-VLmode=Thinking2026.02 | 86.8 | — | — | |
| Doubao-Seed-1.6Model Category=Closed-Source SOTA Models2026.05 | 85.9 | — | — | |
| Qwen3-VL-235B-ThinkingModel Category=Open-Source Large Baselines2026.05 | 85.9 | — | — | |
| TwC (Ours) - Img & TxtCategory=Think with Comic, Note=G-t-R2026.02 | 85.8 | — | — | |
| ERNIE 5.02026.02 | 84.8 | — | — | |
| GLM-4.5VModel Category=Closed-Source SOTA Models2026.05 | 84.6 | — | — | |
| ERNIE 5.0-BaseModel type=pre-trained2026.02 | 84.4 | — | — | |
| Qwen3-VL-32B2025.06 | 83.8 | — | — | |
| Gemini 2.5-Pro2026.02 | 82.7 | — | — | |
| GPT-5tier=High2026.02 | 82.1 | — | — | |
| InternVL3.5-38B2025.06 | 81.9 | — | — | |
| GPT-5-Thinking*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 81.9 | — | — | |
| GPT-5-20250807Model Category=Closed-Source SOTA Models2026.05 | 81.9 | — | — | |
| GenRecal (InternVL3.5-8B)Teacher VLM=Qwen3-VL-32B2025.06 | 81.5 | — | — | |
| GenRecal (InternVL3.5-8B)Teacher VLM=InternVL3.5-38B2025.06 | 81.3 | — | — | |
| Gemini-2.5-Pro*Model Category=Close-source, Reference Source=OpenCompass leaderboard2025.09 | 80.9 | — | — | |
| GenRecal (Qwen3-VL-8B)Teacher VLM=InternVL3.5-38B2025.06 | 80.3 | — | — | |
| GenRecal (Qwen3-VL-8B)Teacher VLM=Qwen3-VL-32B2025.06 | 80.1 | — | — | |
| Kimi-vl-A3B-thinkingModel Category=Open-Source Large Baselines2026.05 | 80.1 | — | — | |
| InternVL3.5-8B2025.06 | 78.4 | — | — | |
| Qwen3-VL-8B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 77.6 | — | — | |
| DEITABase Model=Qwen3-VL-8B-Instruct, Sampling Ratio=Top-20%2026.05 | 77.5 | — | — | |
| Penguin-VLModel Size=8B2026.03 | 77.4 | — | — | |
| OSTBase Model=Qwen3-VL-8B-Instruct, Sampling Ratio=Best-20%2026.05 | 77.4 | — | — | |
| Qwen3-VL-8B2025.06 | 77.2 | — | — | |
| Qwen3-VLModel Size=8B2026.03 | 77.2 | — | — | |
| Qwen3-VL-8B-InstructStrategy=Base2026.05 | 77.2 | — | — | |
| HEEDPipeline Stage=C4+, Post-training=SFT+DPO2026.05 | 77.2 | — | — | |
| VL-Cogito-7B + LEADBackbone=VL-Cogito-7B, Decoding Strategy=LEAD2026.03 | 76.3 | — | — | |
| Qwen3-VL-8B-Instruct + RandomSampling Ratio=20%2026.05 | 76.1 | — | — | |
| TeacherPipeline Stage=C0, Post-training=SFT+DPO2026.05 | 76 | — | — | |
| OSTBase Model=Qwen3-VL-4B-Instruct, Sampling Ratio=Best-20%2026.05 | 75.8 | — | — | |
| VL-Rethinker-7B + LEADBackbone=VL-Rethinker-7B, Decoding Strategy=LEAD2026.03 | 75.6 | — | — | |
| VAPO-Thinker-7BModel Category=Our models, Parameter Scale=7B2025.09 | 75.6 | — | — | |
| RSAPipeline Stage=C3+, Post-training=SFT+DPO2026.05 | 75.4 | — | — | |
| Qwen3-VL-4B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 75.3 | — | — | |
| ThinkLite-VLData Size=11k2026.04 | 75.1 | — | — | |
| DEITABase Model=Qwen3-VL-4B-Instruct, Sampling Ratio=Top-20%2026.05 | 75.1 | — | — | |
| TwC (Ours) - Only ImageCategory=Think with Comic, Note=direct2026.02 | 75 | — | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=InternVL2.5-78B2025.06 | 74.9 | — | — | |
| Vision-R1-7B + LEADBackbone=Vision-R1-7B, Decoding Strategy=LEAD2026.03 | 74.9 | — | — | |
| VL-Rethinker-7BBackbone=VL-Rethinker-7B2026.03 | 74.9 | — | — | |
| VL-RethinkerData Size=39k2026.04 | 74.9 | — | — | |
| V-STARData Size=40k2026.04 | 74.9 | — | — | |
| Qwen3-VL-4B-InstructStrategy=Base2026.05 | 74.9 | — | — | |
| VL-Cogito-7BBackbone=VL-Cogito-7B2026.03 | 74.8 | — | — | |
| VL-CogitoData Size=80k2026.04 | 74.8 | — | — | |
| Qwen2.5-VL-32BParameters=32B2026.02 | 74.7 | — | — | |
| GeoFocusModel Scale=7B2026.02 | 74.3 | — | — | |
| InternVL-3.5Model Size=8B2026.03 | 74.2 | — | — | |
| Qwen3-VL-4B-Instruct + RandomSampling Ratio=20%2026.05 | 74.2 | — | — | |
| RuCL2026.02 | 74.1 | — | — | |
| HSAPipeline Stage=C2+, Post-training=SFT+DPO2026.05 | 74.1 | — | — | |
| MM-Eureka-Qwen-32BActivation Replay=true2025.11 | 74 | — | — | |
| OpenAI-o1Training Paradigm=Closed-source2026.05 | 73.9 | — | — | |
| Vision-R1-7BCategory=Open-source Reasoning Models2026.01 | 73.5 | — | — | |
| Vision-R1-7BBackbone=Vision-R1-7B2026.03 | 73.5 | — | — | |
| Vision-R1-7BModel Category=Open-source, Parameter Scale=7B2025.09 | 73.5 | — | — | |
| Vision-R1Data Size=210k2026.04 | 73.5 | — | — | |
| MM-Eureka-Qwen-7BActivation Replay=true2025.11 | 73.5 | — | — | |
| Vision-R1-7BTraining Paradigm=RL-trained, Parameters=7B2026.05 | 73.5 | — | — | |
| Qwen2.5-VL-7B-Instruct + Faithful-MR1Backbone model=Qwen2.5-VL-7B-Instruct, Training Data=19.2K2026.05 | 73.5 | — | — | |
| KDPipeline Stage=C1+, Post-training=SFT+DPO2026.05 | 73.4 | — | — | |
| ThinkLite-VL-7BParameters=7B2026.02 | 73.3 | — | — | |
| GazeVLM (Ours)Gaze Bias=Enabled, Base Model=Qwen3-VL-4B2026.05 | 73.3 | — | — | |
| VL-Rethinker-7BParameters=7B2026.02 | 73.27 | — | — | |
| Qwen3-VL-4B-Instruct + Full SFTSampling Ratio=100%2026.05 | 73.2 | — | — | |
| Perception-R1-7BTraining Data=1.4K2026.05 | 73.2 | — | — | |
| MM-Eureka-7BParameters=7B2026.02 | 73 | — | — | |
| MM-Eureka-Qwen-7BActivation Replay=false2025.11 | 73 | — | — | |
| Perception-R1-7BParameters=7B2026.02 | 72.8 | — | — | |
| VL-RethinkerParadigm=Tool-use & RL Enhanced Reasoning, Data Size=39K2026.04 | 72.8 | — | — | |
| AutoToolSize=7B2026.05 | 72.8 | — | — | |
| Claude-Sonnet 4.5Category=MLLM, Note=direct2026.02 | 72.5 | — | — | |
| VL-Rethinker-7BActivation Replay=true2025.11 | 72.4 | — | — | |
| InternVL2.5-78B2025.06 | 72.3 | — | — | |
| OpenVLThinkerData Size=59.2k2026.04 | 72.3 | — | — | |
| GRPOModel Scale=7B2026.02 | 72.1 | — | — | |
| MM-Eureka-Qwen-32BActivation Replay=false2025.11 | 72.1 | — | — | |
| VL-Rethinker-7BActivation Replay=false2025.11 | 72 | — | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=InternVL2-76B2025.06 | 71.9 | — | — | |
| DeepEyesV2Tool Use=true, Param Size=7B2026.03 | 71.9 | — | — | |
| MMR1-Math-v0-7BActivation Replay=true2025.11 | 71.9 | — | — | |
| InternVL2.5-38BParameters=38B2026.02 | 71.84 | — | — | |
| InternVL3.5-2BThinking Mode=true2025.12 | 71.8 | — | — | |
| Qwen3-VL-4BTraining Paradigm=Vanilla Open-source, Parameters=4B2026.05 | 71.7 | — | — | |
| InternVL3Tool Use=false, Param Size=8B2026.03 | 71.6 | — | — | |
| Semantic-backTool Use=false, Param Size=7B2026.03 | 71.6 | — | — | |
| InternVL3.5-8BParadigm=Zero-Shot VLMs, Data Size=70K2026.04 | 71.6 | — | — | |
| DeepEyesSize=7B2026.05 | 71.6 | — | — | |
| Gemini-3-ProCategory=MLLM, Note=direct2026.02 | 71.5 | — | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=Qwen2-VL-72B2025.06 | 71.4 | — | — | |
| OpenVLThinker-7BParameters=7B2026.02 | 71.38 | — | — | |
| VRETool Use=false, Param Size=7B2026.03 | 71.2 | — | — | |
| Video-R1-7BCategory=Open-source Reasoning Models2026.01 | 71 | — | — | |
| MMR1-Math-v0-7BActivation Replay=false2025.11 | 71 | — | — | |
| Qwen2.5-VL-7B-Instruct + GRPOBackbone model=Qwen2.5-VL-7B-Instruct, Training Data=19.2K2026.05 | 70.9 | — | — | |
| Vision-R1-7BParameters=7B2026.02 | 70.63 | — | — |