General Visual Reasoning on MMStar (Accuracy)
77.5AccuracyGemini 2.5 Pro
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini 2.5 ProAccess=Close-source2026.02 | 77.5 | — | |
| InternVL3-8BCategory=SFT & Close-source Models, Parameter Size=8B2025.09 | 66.3 | — | |
| SAYO-Qwen-8BBase Model=Qwen3-VL-8B2026.02 | 65.27 | — | |
| GPT4oAccess=Close-source2026.02 | 64.7 | — | |
| GPT-4oReasoning Paradigm=Zero-Shot VLMs2026.01 | 64.7 | — | |
| GPT-4oCategory=SFT & Close-source Models2025.09 | 64.7 | 53.2 | |
| MM-EurekaCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B2025.09 | 64 | 52.52 | |
| NoisyRollout-7BParameters=7B2026.02 | 63.67 | — | |
| SAYO-Qwen-4BBase Model=Qwen3-VL-4B2026.02 | 63.53 | — | |
| Kimi-VL-16BParameters=16B2026.02 | 63.47 | — | |
| VL-RethinkerReasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 63.2 | — | |
| OpenVLThinker-7BParameters=7B2026.02 | 63.07 | — | |
| Semantic-back-7BParameters=7B2026.02 | 63.07 | — | |
| AdaVaR-7BCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B, Reasoning Mode=Adaptive2025.09 | 63 | 55.82 | |
| SAYO-InternVL-8BBase Model=InternVL3.5-8B2026.02 | 62.73 | — | |
| OVR-7BCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B, Reasoning Mode=Text-based2025.09 | 62.7 | 55.76 | |
| Vision-R1Reasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 62.67 | — | |
| Qwen3-VL-8BParameters=8B2026.02 | 62.6 | — | |
| VLAA-Thinker-7BCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B, Reasoning Mode=Text-based2025.09 | 62.6 | 54.2 | |
| ViGoRL2026.02 | 61.53 | — | |
| DeepEyesCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B, Reasoning Mode=Visually-grounded2025.09 | 61.3 | 53.72 | |
| InternVL2.5-Pt + SFTBackbone=InternVL2.5, Training stage=SFT2026.07 | 61.1 | — | |
| Qwen3-VL-4BParameters=4B2026.02 | 60.8 | — | |
| MonetReasoning Paradigm=Latent Reasoning2026.01 | 60.33 | — | |
| Qwen2.5-VL-7BCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B2025.09 | 60.3 | 50.9 | |
| InternVL3.5-8BParameters=8B2026.02 | 60.27 | — | |
| LaserReasoning Paradigm=Latent Reasoning2026.01 | 60.27 | — | |
| InternVL2-Pt + SFTBackbone=InternVL2, Training stage=SFT2026.07 | 60.2 | — | |
| Qwen2.5-VL-7BReasoning Paradigm=Zero-Shot VLMs2026.01 | 59.7 | — | |
| Orsta-7BCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B2025.09 | 59.6 | 52.04 | |
| AdaVaR-3BCategory=Reasoning Models, Base Model=Qwen2.5-VL-3B, Parameter Size=3B, Reasoning Mode=Adaptive2025.09 | 59.3 | 50.84 | |
| Chain-of-FocusCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B2025.09 | 59.3 | 50.94 | |
| LLaVA-OneVisionReasoning Paradigm=Zero-Shot VLMs2026.01 | 59.13 | — | |
| InternVL2-Pt + IRABackbone=InternVL2, Training stage=IRA2026.07 | 58.8 | — | |
| InternVL2.5-Pt + IRABackbone=InternVL2.5, Training stage=IRA2026.07 | 58.8 | — | |
| DeepEyesReasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 58.73 | — | |
| InternVL3.5-30B-A3BParameters=30B-A3B2026.02 | 58.33 | — | |
| LLaVA-OV-Pt + IRABackbone=LLaVA-OV, Training stage=IRA2026.07 | 58.1 | — | |
| LVRReasoning Paradigm=Latent Reasoning2026.01 | 57.93 | — | |
| LLaVA-OV-Pt + SFTBackbone=LLaVA-OV, Training stage=SFT2026.07 | 57.9 | — | |
| R1-Onevision-7BParameters=7B2026.02 | 57 | — | |
| InternVL3.5-14BParameters=14B2026.02 | 55.93 | — | |
| LMM-R1Category=Reasoning Models, Base Model=Qwen2.5-VL-3B, Parameter Size=3B2025.09 | 55 | 48.8 | |
| VLAA-Thinker-3BCategory=Reasoning Models, Base Model=Qwen2.5-VL-3B, Parameter Size=3B, Reasoning Mode=Text-based2025.09 | 54.8 | 46.7 | |
| ViGoRL-7BCategory=Reasoning Models, Base Model=Qwen2.5-VL-7B, Parameter Size=7B2025.09 | 54.3 | 50.53 | |
| InternVL3.5-38BParameters=38B2026.02 | 54.2 | — | |
| Qwen3-VL-30B-A3BParameters=30B-A3B2026.02 | 53.73 | — | |
| InternVL3.5-8BReasoning Paradigm=Zero-Shot VLMs2026.01 | 53.33 | — | |
| Qwen2.5-VL-3BCategory=Reasoning Models, Base Model=Qwen2.5-VL-3B, Parameter Size=3B2025.09 | 53.1 | 45.74 | |
| GRIT-3BCategory=Reasoning Models, Base Model=Qwen2.5-VL-3B, Parameter Size=3B2025.09 | 52.7 | 45.8 | |
| ViGoRL-3BCategory=Reasoning Models, Base Model=Qwen2.5-VL-3B, Parameter Size=3B2025.09 | 51.4 | 45.7 | |
| InternVL2.5-PtBackbone=InternVL2.5, Training stage=Pre-trained2026.07 | 47.3 | — | |
| PAPOReasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 45.8 | — | |
| InternVL2-PtBackbone=InternVL2, Training stage=Pre-trained2026.07 | 43.7 | — | |
| LLaVA-OV-PtBackbone=LLaVA-OV, Training stage=Pre-trained2026.07 | 30.6 | — |