Visual Question Answering on MMStar
91.9AccuracyQwen3.5
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5Parameters=35BA3B, Internal Reasoning (Think mode)=false2026.05 | 91.9 | |
| Gemini 3-Pro2026.02 | 82.96 | |
| GPT-5tier=High2026.02 | 82.1 | |
| SenseNova-U1Parameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 80.92 | |
| Qwen3.5Parameters=9B, Internal Reasoning (Think mode)=false2026.05 | 79.7 | |
| Qwen3-VL 32BModel Source Category=Open-weight VLMs, Model Parameters=32B, Evaluation Mode=Thinking Mode2026.01 | 79.4 | |
| SenseNova-U1Parameters=8B, Internal Reasoning (Think mode)=true2026.05 | 78.27 | |
| Gemini 2.5-Pro2026.02 | 77.5 | |
| Gemma4Parameters=26BA4B, Internal Reasoning (Think mode)=false2026.05 | 76.93 | |
| Qwen3-VLmode=Thinking2026.02 | 76.88 | |
| Gemini-2.5 FlashModel Source Category=Closed-source VLMs, Evaluation Mode=Thinking Mode2026.01 | 76.5 | |
| ERNIE 5.02026.02 | 75.54 | |
| Qwen3-VL 30B-A3BModel Source Category=Open-weight VLMs, Model Parameters=30B-A3B, Evaluation Mode=Thinking Mode2026.01 | 75.5 | |
| Qwen3VLParameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 75.5 | |
| Qwen3-VL 8BModel Source Category=Open-weight VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 75.3 | |
| Qwen3VLParameters=8B, Internal Reasoning (Think mode)=true2026.05 | 75.3 | |
| MMFineReason-8BModel Source Category=Ours, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 75.2 | |
| GPT5 miniModel Source Category=Closed-source VLMs, Evaluation Mode=Thinking Mode2026.01 | 74.1 | |
| HoneyBee 8BModel Source Category=Open-source VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 73.3 | |
| MMFineReason-4BModel Source Category=Ours, Model Parameters=4B, Evaluation Mode=Thinking Mode2026.01 | 72.8 | |
| Qwen3-VL-8BFine-tuning setting=CPO2026.07 | 71.67 | |
| SwimBird2026.02 | 71.2 | |
| Qwen2.5-VL-32B-Instruct2026.02 | 70.3 | |
| MMR1 8BModel Source Category=Open-source VLMs, Model Parameters=8B, Evaluation Mode=Thinking Mode2026.01 | 69.3 | |
| LongCat-NextParameters=68BA3B, Internal Reasoning (Think mode)=false2026.05 | 69.3 | |
| OMR 7BModel Source Category=Open-source VLMs, Model Parameters=7B, Evaluation Mode=Thinking Mode2026.01 | 69 | |
| MMFineReason-2BModel Source Category=Ours, Model Parameters=2B, Evaluation Mode=Thinking Mode2026.01 | 67.7 | |
| SpatialThinker-30BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 66.9 | |
| InternVL3.5Parameters=8B2026.04 | 66.3 | |
| SpatialThinker-7BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 65.9 | |
| InternVL3.5Parameters=4B2026.04 | 65.6 | |
| Claude-3.5-Sonnet-0620Model Category=Proprietary and Open-Source MLLMs2025.11 | 65.1 | |
| BARD-VLParameters=8B, B=42026.04 | 65 | |
| SkiLa2026.02 | 64.8 | |
| Qwen3-VL-8B-Instructreproduced by authors=true2026.02 | 64.7 | |
| GPT-4o-0513Model Category=Proprietary and Open-Source MLLMs2025.11 | 64.7 | |
| Claude-4-Sonnet-0514Model Category=Proprietary and Open-Source MLLMs2025.11 | 64.4 | |
| Qwen3-VL-30BModel Category=Proprietary and Open-Source MLLMs2025.11 | 64.3 | |
| Qwen3-VL-8BFine-tuning setting=LoRA2026.07 | 64.13 | |
| Gemini-ProAccess Type=Closed-source API2024.07 | 64.1 | |
| Qwen2.5-VL-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 63.9 | |
| VLAA-Thinker-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 63.8 | |
| BARD-VLParameters=4B, B=322026.04 | 63.6 | |
| Qwen2.5-VL-7B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 63.4 | |
| Qwen2.5-VL-7B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 63.2 | |
| VanillaSource=Qwen’25, Token Reduction=0%, Backbone=Qwen2.5-VL-7B2026.06 | 62.3 | |
| LLaVA-OneVision2026.02 | 61.9 | |
| Qwen3-VL-8BFine-tuning setting=GSPO2026.07 | 61.67 | |
| GPT-4oModel Category=Proprietary VLMs2024.06 | 61.6 | |
| LLaDA-VParameters=8B2026.04 | 60.4 | |
| Qwen2.5-VL-7B-Instruct2026.02 | 60.3 | |
| IXC-2.5-7BModel Size=7B2024.07 | 59.9 | |
| Qwen3-VLParameters=8B2026.04 | 59.9 | |
| Dream-VLParameters=7B2026.04 | 59.9 | |
| Qwen3-VL-8BFine-tuning setting=GRPO2026.07 | 59.47 | |
| GPT-5-0807Model Category=Proprietary and Open-Source MLLMs2025.11 | 58.9 | |
| SpaceOmModel Category=Proprietary and Open-Source MLLMs2025.11 | 57.7 | |
| SpatialThinker-3BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 57.6 | |
| InternVL-Chat-v1.5Model Category=Open-Source VLMs2024.06 | 57.1 | |
| InternVLModel Size=1.5-26B, Access Type=Open-source2024.07 | 57.1 | |
| GPT-4VAccess Type=Closed-source API2024.07 | 57.1 | |
| Qwen3-VLParameters=4B2026.04 | 56.9 | |
| Qwen2.5-VL-3B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 56.7 | |
| InternLM-XComposer2Model Category=Open-Source VLMs2024.06 | 56.2 | |
| Qwen2.5-VL-3BModel Category=Proprietary and Open-Source MLLMs2025.11 | 55.9 | |
| Qwen3-VL-8BFine-tuning setting=Base2026.07 | 55.87 | |
| SpaceThinkerModel Category=Proprietary and Open-Source MLLMs2025.11 | 54.5 | |
| Qwen2.5-VL-3B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 53.9 | |
| PDrop+RerouteSource=Ours, Token Reduction=66.7%, Backbone=Qwen2.5-VL-7B2026.06 | 53.5 | |
| SDAR-VLParameters=8B2026.04 | 53.3 | |
| DyCo-RLBackbone=Qwen2.5-VL-3B, Zero-shot=true2026.06 | 53.2 | |
| BARD-VLParameters=2B, B=322026.04 | 53.1 | |
| PDropSource=CVPR’25, Token Reduction=66.7%, Backbone=Qwen2.5-VL-7B2026.06 | 52.6 | |
| Qwen3-VL-8BFine-tuning setting=FFT2026.07 | 51.93 | |
| LLaVA-NeXT (Yi-34B)Model Category=Open-Source VLMs2024.06 | 51.6 | |
| GRPO BaselineBackbone=Qwen2.5-VL-3B, Zero-shot=true2026.06 | 51.2 | |
| FastV+RerouteSource=Ours, Token Reduction=66.7%, Backbone=Qwen2.5-VL-7B2026.06 | 50.6 | |
| FastVSource=ECCV’24, Token Reduction=66.7%, Backbone=Qwen2.5-VL-7B2026.06 | 50.3 | |
| GPT-4vModel Category=Proprietary VLMs2024.06 | 49.7 | |
| AvgModel size=7B, Evaluation protocol=0-shot2026.02 | 49.03 | |
| MaD-MixModel size=7B, Evaluation protocol=0-shot2026.02 | 48.79 | |
| UniformModel size=7B, Evaluation protocol=0-shot2026.02 | 48.18 | |
| Dimple-VLParameters=7B2026.04 | 47.7 | |
| PDrop+RerouteSource=Ours, Token Reduction=77.8%, Backbone=Qwen2.5-VL-7B2026.06 | 47.5 | |
| LaviDaParameters=8B2026.04 | 47 | |
| PDropSource=CVPR’25, Token Reduction=77.8%, Backbone=Qwen2.5-VL-7B2026.06 | 46.8 | |
| VanillaSource=Qwen’25, Backbone=Qwen3.5-9B-Hybrid, Average Gated Attention Token Reduction=0%2026.06 | 46.8 | |
| FusedModel size=7B, Evaluation protocol=0-shot2026.02 | 46.55 | |
| Prism Captioner-7BModel Category=Prism Models, Reasoning Module=Llama32024.06 | 45.9 | |
| Prism Captioner-7BModel Category=Prism Models, Reasoning Module=ChatGPT2024.06 | 43.7 | |
| Prism Captioner-2BModel Category=Prism Models, Reasoning Module=ChatGPT2024.06 | 43.3 | |
| FastV+RerouteSource=Ours, Token Reduction=77.8%, Backbone=Qwen2.5-VL-7B2026.06 | 42.4 | |
| Prism Captioner-2BModel Category=Prism Models, Reasoning Module=Llama32024.06 | 42 | |
| LLaVA-InternLM2-20BModel Category=Open-Source VLMs2024.06 | 41.9 | |
| MRoPEPositional Encoding Method=MRoPE, DIPE Enhancement=+DIPE2026.03 | 41.67 | |
| FastVSource=ECCV’24, Token Reduction=77.8%, Backbone=Qwen2.5-VL-7B2026.06 | 41 | |
| PDrop+RerouteSource=Ours, Token Reduction=88.9%, Backbone=Qwen2.5-VL-7B2026.06 | 41 | |
| MRoPE-IPositional Encoding Method=MRoPE-I, DIPE Enhancement=+DIPE2026.03 | 40.73 | |
| Emu2-ChatModel Category=Open-Source VLMs2024.06 | 40.7 | |
| Yi-VL-34BModel Category=Open-Source VLMs2024.06 | 40.5 |