Multi-image visual reasoning on BLINK
69.1AccuracyQwen3-VL
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VLArchitecture Type=Modular, Parameter Scale=8B, Training Stage=Instruct2026.05 | 69.1 | |
| Qwen3-VL-8BModel scale=8B2026.06 | 64.5 | |
| Qwen3-VL-4BModel scale=4B2026.06 | 64.3 | |
| Ouro-Spatial-8BModel scale=8B2026.06 | 64.2 | |
| Ouro-Spatial-4BModel scale=4B2026.06 | 63.1 | |
| NEO-ovArchitecture Type=Native, Parameter Scale=8B, Training Stage=Instruct2026.05 | 62.8 | |
| VAPO2025.09 | 60.1 | |
| InternVL3.5Architecture Type=Modular, Parameter Scale=8B, Training Stage=Instruct2026.05 | 59.5 | |
| GRPO2025.09 | 57 | |
| VideoLLaMA3Architecture Type=Modular, Parameter Scale=8B, Training Stage=Instruct2026.05 | 56.7 | |
| Qwen2.5-VL + S2H-DPOParameter=7B2026.04 | 55.85 | |
| Base model2025.09 | 55.3 | |
| Qwen2.5-VLParameter=7B2026.04 | 54.29 | |
| Qwen3-VL + S2H-DPOParameter=2B2026.04 | 53.92 | |
| NEO-ovArchitecture Type=Native, Parameter Scale=2B, Training Stage=Instruct2026.05 | 53.9 | |
| Qwen3-VLArchitecture Type=Modular, Parameter Scale=2B, Training Stage=Instruct2026.05 | 53.8 | |
| Qwen3-VLParameter=2B2026.04 | 51.61 | |
| InternVL3.5Architecture Type=Modular, Parameter Scale=2B, Training Stage=Instruct2026.05 | 51.3 | |
| GPT-4VParameter=-2026.04 | 51.1 | |
| AVLM-2BModel Scale=2B2026.06 | 46.93 | |
| Idefics2Parameter=8B2026.04 | 45.2 | |
| VideoLLaMA3Architecture Type=Modular, Parameter Scale=2B, Training Stage=Instruct2026.05 | 44.2 | |
| Gemma-4-E2B-itModel Scale=4B2026.06 | 43.93 | |
| LLaVA-v1.5 + S2H-DPOParameter=7B2026.04 | 43.4 | |
| LLaVA-v1.5 + MIA-DPOParameter=7B2026.04 | 42.9 | |
| InstructBLIPParameter=13B2026.04 | 42.2 | |
| CogVLMParameter=17B2026.04 | 41.5 | |
| Qwen2.5-VL + MIA-DPOParameter=7B2026.04 | 41.28 | |
| LLaVA-v1.5 + LLaVA-RLHFParameter=7B2026.04 | 40.8 | |
| LoRATrainable=80M, r=642026.06 | 40.6 | |
| VIFLLM=Vicuna-7B, Resolution=336 × 3362026.04 | 40.5 | |
| LoRATrainable=20M, r=162026.06 | 39.8 | |
| LLaVA-v1.5LLM=Vicuna-7B, Resolution=336 × 3362026.04 | 39.7 | |
| DPVR-PCTrainable=202M, evaluation=3-seed average2026.06 | 39.7 | |
| LLaVA-v1.6Parameter=7B2026.04 | 39.6 | |
| FastVTrainable=0, mode=inference-only2026.06 | 39.6 | |
| DPVR-KVTrainable=202M2026.06 | 39.6 | |
| VideoLLaVAParameter=7B2026.04 | 38.9 | |
| LLaVA-v1.5 + HA-DPOParameter=7B2026.04 | 38.6 | |
| DPVR-LFTrainable=202M, evaluation=3-seed average2026.06 | 38.6 | |
| Vanilla LLaVA-1.5-7BTrainable=02026.06 | 38.5 | |
| IDEFICS-9BLLM=LLaMA 2-7B2026.04 | 38.3 | |
| Mantis-8B-FuyuLLM=Fuyu-8B, Resolution=1024 × 10242026.04 | 38.2 | |
| LLaVA-v1.5Parameter=7B2026.04 | 37.1 | |
| FuyuParameter=8B2026.04 | 36.6 | |
| Qwen2.5-Omni-3BModel Scale=3B2026.06 | 36.39 | |
| Emu2-ChatParameter=37B2026.04 | 36.2 | |
| Qwen-VL-ChatParameter=7B2026.04 | 31.2 | |
| Qwen-VL-ChatLLM=Qwen-7B, Resolution=448 × 4482026.04 | 28.2 | |
| Qwen-VLLLM=Qwen-7B, Resolution=448 × 4482026.04 | 27.9 | |
| LLaVA-v1.5 + POVIDParameter=7B2026.04 | 19.9 |