Visual Question Answering on RealworldQA
80.2AccuracyL2-VMAS
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| L2-VMASBackbone=Qwen3-VL-8B-Thinking2026.01 | 80.2 | 3,976 | |
| Qwen3-VL-32B-Instruct + SpatialBoostEncoder Enhancement=SpatialBoost2026.03 | 79.6 | — | |
| Qwen3-VL-32B-Instruct2026.03 | 79 | — | |
| GPT-5 miniVersion=high2026.05 | 79 | — | |
| GPT-5-0807Model Category=Proprietary and Open-Source MLLMs2025.11 | 78.7 | — | |
| RLVRTraining Source=ViRL2026.05 | 77.9 | — | |
| L2-VMASBackbone=Qwen3-VL-8B-Instruct2026.01 | 77.6 | 2,101 | |
| RLR³Training Source=OpenMMR2026.05 | 77.6 | — | |
| Official thinkingSource=Qwen3-VL2026.05 | 77.4 | — | |
| RLR³Training Source=ViRL2026.05 | 76.7 | — | |
| RLVRTraining Source=OpenMMR2026.05 | 76.7 | — | |
| RLVRTraining Source=DeepVision2026.05 | 76.6 | — | |
| L2-VMASBackbone=LLaVA-OV-1.5-8B2026.01 | 76.2 | 2,135 | |
| VMASBackbone=Qwen3-VL-8B-Thinking2026.01 | 76.2 | 7,026 | |
| InternVL3-38B + SpatialBoostEncoder Enhancement=SpatialBoost2026.03 | 75.9 | — | |
| Penguin-VLModel Size=8B2026.03 | 75.8 | — | |
| InternVL3-38B2026.03 | 75.6 | — | |
| RLR³Training Source=DeepVision2026.05 | 75.6 | — | |
| GPT-4o-0513Model Category=Proprietary and Open-Source MLLMs2025.11 | 75.4 | — | |
| SpatialThinker-30BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 74.9 | — | |
| L2-VMASBackbone=InternVL-3.5-8B2026.01 | 74.8 | 2,534 | |
| VMASBackbone=Qwen3-VL-8B-Instruct2026.01 | 74.7 | 2,769 | |
| Base instruct2026.05 | 74.4 | — | |
| FULLModel=Qwen3-VL (4B), Selection Ratio=100%2026.05 | 73.7 | — | |
| Official instructSource=Qwen3-VL2026.05 | 73.7 | — | |
| GPT-5 miniVersion=minimal2026.05 | 73.3 | — | |
| Honeybee-Remake-SEED-200KModel=Qwen3-VL (4B), Selection Ratio=20%2026.05 | 73.2 | — | |
| SingleBackbone=Qwen3-VL-8B-Thinking2026.01 | 72.9 | 730 | |
| VMASBackbone=LLaVA-OV-1.5-8B2026.01 | 72.8 | 2,891 | |
| VMASBackbone=InternVL-3.5-8B2026.01 | 72.7 | 3,360 | |
| Qwen3-VL-4BModel=Qwen3-VL, Parameter Count=4B2026.04 | 71.9 | — | |
| HeadLensBackbone=Qwen3-VL-8B2026.03 | 71.83 | — | |
| RANDOMModel=Qwen3-VL (4B), Selection Ratio=20%2026.05 | 71.8 | — | |
| Qwen3-VL 4B + 2B-RMain Model=4B, Source Model=2B-R, Strategy=Reasoning Transfer2026.04 | 71.6 | — | |
| Qwen3-VLModel Size=8B2026.03 | 71.5 | — | |
| BASEModel=LLaVA-OneVision-1.5 (4B), Selection Ratio=0%2026.05 | 71.4 | — | |
| Qwen3-VL 4B-SModel Scale=4B, Strategy=Simple2026.04 | 71.2 | — | |
| BASEModel=Qwen3-VL (4B), Selection Ratio=0%2026.05 | 71.2 | — | |
| AutoVModel=Qwen2.5-VL 7B2025.06 | 71.1 | — | |
| SingleBackbone=Qwen3-VL-8B-Instruct2026.01 | 71 | 402 | |
| RISEBackbone=Qwen3-VL-8B-Instruct, Self-evolving steps=402026.05 | 70.85 | — | |
| RISEBackbone=Qwen3-VL-8B-Instruct, Self-evolving steps=602026.05 | 70.33 | — | |
| L2-VMASBackbone=GLM-4.1V-9B-Thinking2026.01 | 70.3 | 4,314 | |
| Qwen3-VL 8B-RModel Scale=8B, Strategy=Reasoning2026.04 | 70.3 | — | |
| Penguin-VLParameters=2B2026.03 | 70.2 | — | |
| Qwen3-VL 8B + 4B-RMain Model=8B, Source Model=4B-R, Strategy=Reasoning Transfer2026.04 | 70.2 | — | |
| DefenderIteration=32026.01 | 70.07 | — | |
| BaselineBackbone=Qwen3-VL-8B2026.03 | 70.05 | — | |
| Qwen3-VL 4B-RModel Scale=4B, Strategy=Reasoning2026.04 | 69.9 | — | |
| OriginalCompression Ratio (CR)=None2026.02 | 69.7 | — | |
| Qwen 3 VL 8BCR=uncompressed2026.02 | 69.67 | — | |
| LLaVA-OneVision-7BModel=LLaVA-OneVision, Parameter Count=7B2026.04 | 69.54 | — | |
| FULLModel=LLaVA-OneVision-1.5 (4B), Selection Ratio=100%2026.05 | 69.4 | — | |
| Base (M_def^(0)) + Clean DataData=Cleaned2026.01 | 69.28 | — | |
| DefenderIteration=12026.01 | 69.28 | — | |
| DefenderIteration=22026.01 | 69.28 | — | |
| APIModel=Qwen2.5-VL 7B2025.06 | 69.2 | — | |
| SpatialThinker-7BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 69.2 | — | |
| SingleBackbone=GLM-4.1V-9B-Thinking2026.01 | 69 | 708 | |
| RANDOMModel=LLaVA-OneVision-1.5 (4B), Selection Ratio=20%2026.05 | 69 | — | |
| AutoVModel=LLaVA-OneVision 7B2025.06 | 68.9 | — | |
| Qwen3-VL 8B + 2B-RMain Model=8B, Source Model=2B-R, Strategy=Reasoning Transfer2026.04 | 68.9 | — | |
| SpaRE-7BModel size=7B2025.04 | 68.8 | — | |
| Honeybee-Remake-SEED-200KModel=LLaVA-OneVision-1.5 (4B), Selection Ratio=20%2026.05 | 68.8 | — | |
| Qwen3-VL-2B-Instruct + MoDA (ours)LLM=Qwen3-2B2025.06 | 68.8 | — | |
| BaseModel=Qwen2.5-VL 7B2025.06 | 68.5 | — | |
| RISEBackbone=Qwen3-VL-8B-Instruct, Self-evolving steps=202026.05 | 68.5 | — | |
| Qwen2.5-VL-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 68.4 | — | |
| SingleBackbone=LLaVA-OV-1.5-8B2026.01 | 68.3 | 445 | |
| WeMMModel Size=8B, Access Type=Open-source2024.07 | 68.1 | — | |
| SR-3DModel Version=base, Inference Mode=base2026.06 | 68.1 | — | |
| GPT-4VAccess Type=Closed-source API2024.07 | 68 | — | |
| IXC-2.5-7BModel Size=7B2024.07 | 67.8 | — | |
| Base (M_def^(0))2026.01 | 67.71 | — | |
| Qwen2VL-7BModel size=7B2025.04 | 67.7 | — | |
| Qwen3-VL 8B-SModel Scale=8B, Strategy=Simple2026.04 | 67.6 | — | |
| InternVL-3.5Model Size=8B2026.03 | 67.5 | — | |
| Base ModelBackbone=Qwen3-VL-8B-Instruct, Self-evolving steps=02026.05 | 67.45 | — | |
| SingleBackbone=InternVL-3.5-8B2026.01 | 67.4 | 523 | |
| VideoLLaMA3Architecture Category=Token Insertion – Proprietary2025.12 | 67.3 | — | |
| InfiniteVL-4BModel=InfiniteVL, Parameter Count=4B, Result Source=InfiniteVL (Tao et al., 2025) paper2026.04 | 67.3 | — | |
| VMASBackbone=GLM-4.1V-9B-Thinking2026.01 | 67.2 | 7,446 | |
| SpB2.0-VL-5BModel=SpB2.0-VL, Parameter Count=5B2026.04 | 67.19 | — | |
| Qwen2.5-VL-7B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 66.6 | — | |
| APIModel=LLaVA-OneVision 7B2025.06 | 66.4 | — | |
| VLAA-Thinker-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 66.4 | — | |
| SpatialThinker-3BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 66.3 | — | |
| BaseModel=LLaVA-OneVision 7B2025.06 | 66 | — | |
| Qwen2.5-VL-7B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 65.4 | — | |
| Qwen2.5-VL-3BModel=Qwen2.5-VL, Parameter Count=3B2026.04 | 65.1 | — | |
| LLaDA-FastVSettings=K=15, P=50, Avg (%)=98.7%, FLOPs (Rem./Red.)=51% (49%↓)2026.01 | 64.84 | — | |
| Qwen3-VL-30BModel Category=Proprietary and Open-Source MLLMs2025.11 | 64.8 | — | |
| Qwen2.5-VL-3B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 64.8 | — | |
| Qwen3-VL-2B-InstructLLM=Qwen3-2B2025.06 | 64.7 | — | |
| SR-ReaL-directModel Version=Full, Inference Mode=direct2026.06 | 64.6 | — | |
| InternVL2-8BModel size=8B2025.04 | 64.4 | — | |
| Qwen2.5-VL-3B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 64.4 | — | |
| Cambrian 8B# Vis tok.=5762024.12 | 64.2 | — | |
| Florence-VL 8B# Vis tok.=5762024.12 | 64.2 | — | |
| HeadLensBackbone=Qwen2.5-VL-7B2026.03 | 64.14 | — |