Visual Reasoning on BLINK
85.2AccuracyQwen3-VL-8B-Inst.
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL-8B-Inst.Model Type=Open-source general-purpose MLLM2026.03 | 85.2 | |
| Vlaser-8BModel Type=Embodied Brain MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 84.9 | |
| RoboBrain2.5-8BModel Type=Embodied Brain MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 84.3 | |
| InternVL3.5-8BModel Type=Open-source general-purpose MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 84.1 | |
| ACE-Brain-0-8BModel Type=Embodied Brain MLLM2026.03 | 83.9 | |
| InternVL3-8BModel Type=Open-source general-purpose MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 82.9 | |
| Qwen2.5-VL-7B-Inst.Model Type=Open-source general-purpose MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 82.5 | |
| Gemini-2.5-ProModel Type=Closed-source MLLM2026.03 | 81.8 | |
| RoboBrain2.0-7BModel Type=Embodied Brain MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 81.4 | |
| GPT-5-miniGrounded=false2025.12 | 81 | |
| VeBrain-7BModel Type=Embodied Brain MLLM2026.03 | 79.7 | |
| Claude-4-SonnetModel Type=Closed-source MLLM2026.03 | 78.1 | |
| GPT-4oModel Type=Closed-source MLLM2026.03 | 77.9 | |
| GPT-5-miniGrounded=true2025.12 | 76 | |
| Qwen3-VL-2B-Inst.Model Type=Open-source general-purpose MLLM2026.03 | 74.9 | |
| Gemini-2.5-ProData Scale=-, Training Phases=-2026.01 | 70.6 | |
| GRITBase Model=Qwen2.5-VL-3B2025.12 | 70.3 | |
| VALOR2025.12 | 69.2 | |
| VALORBase Model=Qwen3-8B2025.12 | 69.2 | |
| VALORModel Category=Open-Source, Tool Use=true2025.12 | 69.2 | |
| ViGoRLBase Model=Qwen2.5-VL-7B2025.12 | 68.4 | |
| GPT-4oReasoning Paradigm=Zero-Shot VLMs2026.01 | 68 | |
| GPT-4omodel_category=Proprietary Model2025.09 | 68 | |
| VALOR-RLBase Model=Qwen3-8B2025.12 | 67.3 | |
| VALOR-RLModel Category=Open-Source, Tool Use=true2025.12 | 67.3 | |
| Qwen3-VLScale=4B2026.05 | 66.6 | |
| o4-miniModel Category=Proprietary, Tool Use=true2025.12 | 66.5 | |
| GPT-4oData Scale=-, Training Phases=-2026.01 | 65.9 | |
| Claude-3.5-HaikuModel Category=Proprietary, Tool Use=true2025.12 | 64.6 | |
| GPT-4oModel Category=Proprietary, Tool Use=true2025.12 | 64.2 | |
| Qwen3-8BModel Category=Open-Source, Tool Use=true2025.12 | 63.9 | |
| GPT-4oModel Type=Proprietary2026.05 | 63.55 | |
| Gemini2.0-Flashmodel_category=Proprietary Model2025.09 | 63.5 | |
| Qwen3.5-VLScale=4B2026.05 | 63.5 | |
| SmoothOp-7BData Scale=50K, Training Phases=1 (R)2026.01 | 62.6 | |
| VST-7B-SFTData Scale=4.1M, Training Phases=1 (S)2026.01 | 62.1 | |
| Qwen3.5-VLScale=2B2026.05 | 62 | |
| Gemini-2.5-FlashModel Category=Proprietary, Tool Use=true2025.12 | 61.5 | |
| RIS+VLPOBackbone=Qwen2.5-VL-7B, Latent Tokens=5, Optimization=Visual-latent Policy Optimization (VLPO)2026.05 | 60.95 | |
| RISBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 60.6 | |
| VST-3B-SFTData Scale=4.1M, Training Phases=1 (S)2026.01 | 59.1 | |
| InternVL3.5Scale=4B2026.05 | 57.8 | |
| CoVTReasoning Paradigm=Latent-Space Reasoning, Reasoning Space=Latent2026.07 | 57.49 | |
| ProLaViT (Base)Reasoning Paradigm=Progressive Latent Visual Thought, Reasoning Space=Latent2026.07 | 57.49 | |
| Gemma-3-12BModel Category=Open-Source, Tool Use=true2025.12 | 57.4 | |
| CoVTBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 57.4 | |
| ProLaViT (Full)Reasoning Paradigm=Progressive Latent Visual Thought, Reasoning Space=Latent, Training Strategy=Distance-Weighted Diversity Loss2026.07 | 57.23 | |
| VisionZero-Qwen-7B (Real-World)training_context=Real-World, base_model=Qwen2.5-VL-7B2025.09 | 57.2 | |
| LaserReasoning Paradigm=Latent Reasoning2026.01 | 56.92 | |
| SpatialLadder-3BData Scale=26K, Training Phases=3 (S+R)2026.01 | 56.9 | |
| SmoothOp-3BData Scale=50K, Training Phases=1 (R)2026.01 | 56.9 | |
| Pelican-VL-7BModel Type=Embodied Brain MLLM2026.03 | 56.8 | |
| VisionZero-Qwen-7B (Chart)training_context=Chart, base_model=Qwen2.5-VL-7B2025.09 | 56.8 | |
| LVRBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 56.79 | |
| MonetBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 56.7 | |
| CoT SFTReasoning Paradigm=Text-Space Reasoning, Reasoning Space=Text2026.07 | 56.54 | |
| Qwen2.5-VL-7BData Scale=-, Training Phases=-2026.01 | 56.4 | |
| ProLaViT + L_div^marginReasoning Paradigm=Progressive Latent Visual Thought, Reasoning Space=Latent2026.07 | 56.33 | |
| One-step Latent Pred.Reasoning Paradigm=Latent-Space Reasoning, Reasoning Space=Latent2026.07 | 56.28 | |
| Qwen2.5-VL-7B+GLSDBackbone=Qwen2.5-VL-7B, Training=Grounded Latent Supervision Dataset (GLSD)2026.05 | 56.25 | |
| ViLaSR-7BData Scale=81K, Training Phases=2 (S+R)2026.01 | 56.2 | |
| VisionZero-Qwen-7B (CLEVR)training_context=CLEVR, base_model=Qwen2.5-VL-7B2025.09 | 56 | |
| Qwen3-VLScale=2B2026.05 | 55.7 | |
| ViGaL-Snake+Rotationbase_model=Qwen2.5-VL-7B2025.09 | 55.6 | |
| SFTReasoning Paradigm=Text-Space Reasoning, Reasoning Space=Text2026.07 | 55.6 | |
| VL-RethinkerReasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 55.55 | |
| InternVL3-8BData Scale=-, Training Phases=-2026.01 | 55.5 | |
| SpaceR-7BData Scale=151K, Training Phases=2 (S+R)2026.01 | 55.4 | |
| Qwen2.5-VL-7Bbase_model=Qwen2.5-VL-7B2025.09 | 55.2 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 54.94 | |
| InternVL3.5-8BReasoning Paradigm=Zero-Shot VLMs2026.01 | 54.81 | |
| Qwen2.5-VL-InstructReasoning Paradigm=Base Model, Reasoning Space=N/A2026.07 | 54.49 | |
| Llama-3.2-11BModel Category=Open-Source, Tool Use=true2025.12 | 54.3 | |
| MM-Eureka-Qwen-7Bbase_model=Qwen2.5-VL-7B2025.09 | 54 | |
| Qwen2.5-VL-7BReasoning Paradigm=Zero-Shot VLMs2026.01 | 53.6 | |
| LVRReasoning Paradigm=Latent Reasoning2026.01 | 53.6 | |
| LVRReasoning Paradigm=Latent-Space Reasoning, Reasoning Space=Latent2026.07 | 53.02 | |
| VLAA-Thinker-7Bbase_model=Qwen2.5-VL-7B2025.09 | 53 | |
| Vision-R1Reasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 52.71 | |
| LLaVALLM=Qwen2-7B2025.12 | 52.7 | |
| PAPOReasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 52.66 | |
| InternVL3.5Scale=2B2026.05 | 52.6 | |
| Gemini-2.0-FlashModel Category=Proprietary, Tool Use=true2025.12 | 51.7 | |
| VISTA-SFTTraining Dataset=SLAKE, Model=Qwen2.5-3B-VL, Training Protocol=single-round self-improvement SFT2026.05 | 51.23 | |
| DeepEyesReasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 51.08 | |
| MonetReasoning Paradigm=Latent Reasoning2026.01 | 50.71 | |
| OpenVLThinker-7Bbase_model=Qwen2.5-VL-7B2025.09 | 49.9 | |
| STaRTraining Dataset=SLAKE, Model=Qwen2.5-3B-VL, Training Protocol=single-round self-improvement SFT2026.05 | 49.78 | |
| Ours w/o MaskingLLM=Qwen2-7B2025.12 | 49.6 | |
| SFT-SeedTraining Dataset=SLAKE, Model=Qwen2.5-3B-VL, Training Protocol=single-round self-improvement SFT2026.05 | 49.35 | |
| LLaVA-OneVisionReasoning Paradigm=Zero-Shot VLMs2026.01 | 49.34 | |
| VIRALLLM=Qwen2-7B2025.12 | 49.3 | |
| JARVISLLM=Qwen2-7B2025.12 | 49.3 | |
| R1-OneVision-7Bbase_model=Qwen2.5-VL-7B2025.09 | 48.7 | |
| JARVISLLM=Vicuna-7B2025.12 | 48.3 | |
| LLava-OneVision-7BData Scale=-, Training Phases=-2026.01 | 48.2 | |
| ReSTEMTraining Dataset=SLAKE, Model=Qwen2.5-3B-VL, Training Protocol=single-round self-improvement SFT2026.05 | 47.62 | |
| Qwen2.5-VL-3BData Scale=-, Training Phases=-2026.01 | 47.6 | |
| LLaVALLM=Vicuna-7B2025.12 | 46.8 | |
| DatologyScale=4B2026.05 | 45.9 |