Multi-image Reasoning on MuirBench
77.2AccuracyL2-VMAS
Evaluation Results
| Method | Links | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| L2-VMASBackbone=Qwen3-VL-8B, Mode=Thinking2026.01 | 77.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SingleBackbone=Qwen3-VL-8B, Mode=Thinking2026.01 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisPlay-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 75 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVE (Ours-8B-iter4)Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=42026.04 | 74.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-InstructModel Category=Open-Source MLLMs, Scale=8B2026.04 | 73.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Jigsaw-R1-8BModel Category=Template-based Self-evolution Methods, Scale=8B2026.04 | 73.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VMASBackbone=Qwen3-VL-8B, Mode=Thinking2026.01 | 73 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM-Zero-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 72.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DAPO (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DAPO2026.01 | 71.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Two-stage RL (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=Two-stage RL (DPS and annealing)2026.01 | 71 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oModel Category=Closed-Source MLLMs2026.01 | 68 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o2026.01 | 68 | 44.51 | 56.12 | 51.28 | 49.15 | 88.69 | 60.29 | 56 | 86.85 | 23.44 | 71.51 | 36.9 | 80.14 | — | |
| GPT-4o-20240513Model Category=Closed-Source MLLMs2026.04 | 68 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DPS (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DPS2026.01 | 67.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| L2-VMASBackbone=Qwen3-VL-8B, Mode=Instruct2026.01 | 67.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CGC-8BBase MLLM=Qwen3-VL2026.04 | 66.42 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-5 nano (high)Model Category=Closed-Source MLLMs, Scale=high2026.04 | 65.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VMASBackbone=Qwen3-VL-8B, Mode=Instruct2026.01 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VLArchitecture Type=Modular, Parameter Scale=8B, Training Stage=Instruct2026.05 | 64.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SingleBackbone=Qwen3-VL-8B, Mode=Instruct2026.01 | 64.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B2026.04 | 63.54 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4VModel Size=Proprietary2026.01 | 62.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VAPO2025.09 | 62.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VLAA-Thinker-7BBase MLLM=Qwen2.5-VL2026.04 | 61 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CGC-7BBase MLLM=Qwen2.5-VL2026.04 | 60.88 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM-Eureka-7BBase MLLM=Qwen2.5-VL2026.04 | 60.57 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MiCo-7BBase MLLM=Qwen2.5-VL2026.04 | 60.53 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO2025.09 | 60.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8BModel scale=8B2026.06 | 60 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NoisyRollout-7BBase MLLM=Qwen2.5-VL2026.04 | 59.61 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-4BModel scale=4B2026.06 | 59 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ouro-Spatial-8BModel scale=8B2026.06 | 58.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B2026.04 | 58.43 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GeM-VG2026.01 | 58.2 | 54.88 | 50 | 64.1 | 60.68 | 68.34 | 57.35 | 50 | 59.05 | 50 | 66.13 | 48.81 | 50 | — | |
| NEO-ovArchitecture Type=Native, Parameter Scale=8B, Training Stage=Instruct2026.05 | 58.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ouro-Spatial-4BModel scale=4B2026.06 | 58.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Base model2025.09 | 58.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLModel Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 57.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PulseFocus (InternVL3.5)Params=8B, Configuration=Gating2026.03 | 57.88 | — | — | — | — | — | — | — | — | — | — | — | — | 1.07 | |
| Migician2026.01 | 57.81 | 54.27 | 50 | 58.97 | 55.13 | 65.83 | 52.65 | 50 | 63.58 | 50 | 74.19 | 46.43 | 50 | — | |
| ThinkLite-VL-7BBase MLLM=Qwen2.5-VL2026.04 | 57.62 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-5 mini (minimal)Model Category=Closed-Source MLLMs, Scale=minimal2026.04 | 57.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5Params=8B, Configuration=Baseline2026.03 | 56.81 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NEO-ovArchitecture Type=Native, Parameter Scale=2B, Training Stage=Instruct2026.05 | 56.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VideoRFTModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 56.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PulseFocus (Qwen3-VL)Params=4B, Configuration=Budget2026.03 | 56.38 | — | — | — | — | — | — | — | — | — | — | — | — | 0.82 | |
| TW-GRPOModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 55.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5Architecture Type=Modular, Parameter Scale=8B, Training Stage=Instruct2026.05 | 55.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VLParams=4B, Configuration=Baseline2026.03 | 55.56 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3-8B2026.04 | 55.04 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OneVision-72BModel Category=Open-Source MLLMs, Scale=72B2026.04 | 54.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Migician-7BBase MLLM=Qwen2-VL2026.04 | 53.69 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3-9BModel Category=Open-Source MLLMs, Scale=9B2026.04 | 51.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ours (masked) (LLaVA-OV-7B)Model Size=7B, Backbone=LLaVA-OV, Attention Masking=True2026.01 | 51.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL2-Llama3-76BModel Size=76B, Backbone=Llama32026.01 | 51.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL2.5Model Category=Open-Source General MLLMs, Parameter Scale=8B2026.01 | 51.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini Pro2026.01 | 49.35 | 35.98 | 41.33 | 47.44 | 28.63 | 64.82 | 45.29 | 48 | 66.59 | 12.5 | 59.14 | 28.57 | 43.84 | — | |
| Qwen3-VLArchitecture Type=Modular, Parameter Scale=2B, Training Stage=Instruct2026.05 | 47.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AVLM-2BModel Scale=2B2026.06 | 46.88 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Full CacheBackbone=Qwen2.5-VL-7B-Instruct2025.11 | 45.54 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CcDPOModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 44.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UniVG-R12026.01 | 44.77 | 39.63 | 46.94 | 44.87 | 42.74 | 52.26 | 44.12 | 27 | 55.17 | 12.5 | 67.74 | 30.95 | 24.32 | — | |
| Mantis-Idefics2Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 44.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VISCModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 44.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Mantis2026.01 | 44.5 | 33.54 | 48.47 | 38.46 | 38.46 | 67.59 | 28.82 | 26 | 53.88 | 18.75 | 56.99 | 26.19 | 35.62 | — | |
| SnapKVretention ratio ρ=0.05, Backbone=Qwen2.5-VL-7B-Instruct2025.11 | 44.42 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlashCacheretention ratio ρ=0.1, Backbone=Qwen2.5-VL-7B-Instruct2025.11 | 44.42 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlashCacheretention ratio ρ=0.05, Backbone=Qwen2.5-VL-7B-Instruct2025.11 | 44.42 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5Architecture Type=Modular, Parameter Scale=2B, Training Stage=Instruct2026.05 | 44 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2VL-7BModel Size=7B2026.01 | 43 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OneVisionModel Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 41.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OV-7BModel Size=7B2026.01 | 41.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BParameters=7B2026.01 | 39.88 | 38.41 | 49.49 | 42.31 | 40.6 | 39.7 | 31.76 | 26 | 51.51 | 12.5 | 68.28 | 32.14 | 19.18 | — | |
| Qwen2-VL-7B2026.04 | 39.88 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ours (LLaVA-OV-1.5B)Model Size=1.5B, Backbone=LLaVA-OV, Attention Masking=False2026.01 | 39.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL2-8BModel Size=8B2026.01 | 37.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| StreamingLLMretention ratio ρ=0.1, Backbone=Qwen2.5-VL-7B-Instruct2025.11 | 37.58 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| StreamingLLMretention ratio ρ=0.05, Backbone=Qwen2.5-VL-7B-Instruct2025.11 | 37.58 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Omni-3BModel Scale=3B2026.06 | 35.15 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| mPLUG-Owl3Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 34 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ours (LLaVA-OV-0.5B)Model Size=0.5B, Backbone=LLaVA-OV, Attention Masking=False2026.01 | 33.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VILA1.5Model Category=Open-Source General MLLMs, Parameter Scale=8B2026.01 | 33.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ours (masked) (LLaVA-OV-0.5B)Model Size=0.5B, Backbone=LLaVA-OV, Attention Masking=True2026.01 | 32.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ours (masked) (LLaVA-OV-1.5B)Model Size=1.5B, Backbone=LLaVA-OV, Attention Masking=True2026.01 | 32.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3 + FOCUSModel Size=≥7B2025.08 | 31.88 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3Model Size=≥7B2025.08 | 31.62 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL + FOCUSModel Size=<7B2025.08 | 31.31 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OV-1.5BModel Size=1.5B2026.01 | 31.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-NeXT-InterleaveModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 31.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Avg + FOCUSModel Size=≥7B2025.08 | 30.47 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLModel Size=<7B2025.08 | 30.38 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AvgModel Size=≥7B2025.08 | 30.09 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLModel Size=≥7B2025.08 | 29.92 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OV + FOCUSModel Size=≥7B2025.08 | 29.81 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL + FOCUSModel Size=≥7B2025.08 | 29.73 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OVModel Size=≥7B2025.08 | 28.73 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3 + FOCUSModel Size=<7B2025.08 | 28.58 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Avg + FOCUSModel Size=<7B2025.08 | 28.17 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3Model Size=<7B2025.08 | 28.04 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA 1.6Model Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 27.4 | — | — | — | — | — | — | — | — | — | — | — | — | — |