Visual Question Answering on MMVP
80.33AccuracyDART
Evaluation Results
| Method | Links | |
|---|---|---|
| DARTAgent=Qwen, GLM, InternVL2026.05 | 80.33 | |
| EAGLEAgent=Qwen, GLM, InternVL2026.05 | 80.33 | |
| Qwen2.5-VL-7B +PCBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 79.3 | |
| Debate(Judge)Agent=Qwen, GLM, InternVL2026.05 | 78 | |
| Qwen2.5-VL-7BBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=No2026.03 | 77.7 | |
| Debate(Vote)Agent=Qwen, GLM, InternVL2026.05 | 77.67 | |
| ReConcileAgent=Qwen, GLM, InternVL2026.05 | 77 | |
| Zero-shot CoTAgent=GLM2026.05 | 76 | |
| Zero-shot CoTAgent=Qwen2026.05 | 75.33 | |
| Self-ConsistencyAgent=Qwen2026.05 | 75.33 | |
| TreeVGR-7BParameters=7B2025.07 | 75.3 | |
| MiMo-VL-7B +PCBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 74.3 | |
| HyLaR-7BModel Type=Visual-Latent Model, Number of Parameters=7B2026.04 | 73.67 | |
| LLaVA-OneVisionModel Type=Open-Source Model2026.04 | 73 | |
| MiMo-VL-7BBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=No2026.03 | 72.7 | |
| ZoomEyeModel Type=Thinking-with-Images Agent Model2026.04 | 72.67 | |
| LaserModel Type=Visual-Latent Model2026.04 | 72 | |
| HyLaR-SFTModel Type=Visual-Latent Model, Training Protocol=SFT2026.04 | 71 | |
| Qwen2.5-VL-3BBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=No2026.03 | 70.7 | |
| Qwen2.5-VL-3B +PCBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=Yes2026.03 | 70.3 | |
| DeepEyesModel Type=Thinking-with-Images Agent Model2026.04 | 70 | |
| Zero-shot CoTAgent=InternVL2026.05 | 69 | |
| MonetModel Type=Visual-Latent Model2026.04 | 68 | |
| Qwen2.5-VL-7BParameters=7B2025.07 | 66.7 | |
| Qwen2.5-VL-72BParameters=72B2025.07 | 66.7 | |
| Qwen2.5-VL-7BModel Type=Open-Source Model, Number of Parameters=7B2026.04 | 65.67 | |
| LVRModel Type=Visual-Latent Model2026.04 | 64 | |
| Vanilla RoPEPositional Encoding Method=Vanilla RoPE, DIPE Enhancement=+DIPE2026.03 | 61.33 | |
| MRoPE-IPositional Encoding Method=MRoPE-I, DIPE Enhancement=+DIPE2026.03 | 61.33 | |
| MRoPE-IPositional Encoding Method=MRoPE-I, DIPE Enhancement=Base2026.03 | 60.67 | |
| Self-RefineAgent=Qwen2026.05 | 60.33 | |
| MRoPEPositional Encoding Method=MRoPE, DIPE Enhancement=+DIPE2026.03 | 60 | |
| Vanilla RoPEPositional Encoding Method=Vanilla RoPE, DIPE Enhancement=Base2026.03 | 58.67 | |
| InternVL3.5-8BModel Type=Open-Source Model, Number of Parameters=8B2026.04 | 57.67 | |
| MRoPEPositional Encoding Method=MRoPE, DIPE Enhancement=Base2026.03 | 56.67 | |
| LLaVA-OV-7BKV Size=1962026.05 | 54 | |
| E2EVLM=Qwen2.5-VL, Setting=RL2025.12 | 53.33 | |
| E2EVLM=Qwen2.5-VL, Setting=base + CoT2025.12 | 52.67 | |
| VISTAVLM=Llama3.2-Vision, Setting=RL2025.12 | 52.67 | |
| E2EVLM=Qwen2.5-VL, Setting=base2025.12 | 51.33 | |
| AttWarp-ChainsBase MLLM=LLaVA2025.10 | 51 | |
| AttWarpBase MLLM=LLaVA2025.10 | 50.7 | |
| E2EVLM=Qwen2.5-VL, Setting=SFT2025.12 | 50.67 | |
| VISTAVLM=Qwen2.5-VL, Setting=RL2025.12 | 50 | |
| AttWarp-DistillBase MLLM=LLaVA2025.10 | 49.3 | |
| LLaVA-Next-Qwen2-7BKV Size=7842026.05 | 49.3 | |
| Base MLLMBase MLLM=LLaVA2025.10 | 48.3 | |
| E2EVLM=Llama3.2-Vision, Setting=base + CoT2025.12 | 48 | |
| Cambrian-4B + TokenCompre.KV Size=1442026.05 | 47.3 | |
| Cambrian-4B + TokenCompre. + Gaze attentionKV Size=54+42026.05 | 47.3 | |
| VISTAVLM=Qwen2.5-VL, Setting=base2025.12 | 46.67 | |
| Cambrian-4B + Gaze attentionKV Size=144+42026.05 | 45.7 | |
| E2EVLM=Llama3.2-Vision, Setting=base2025.12 | 45.33 | |
| LLaVA-OV-7B + HERMESKV Size=1002026.05 | 45.3 | |
| Cambrian-4BKV Size=5762026.05 | 45.3 | |
| Cambrian-4B + Gaze attentionKV Size=288+42026.05 | 45.3 | |
| Cambrian-4B + Gaze attentionKV Size=72+42026.05 | 43.3 | |
| Qwen2.5-VL-3BKV Size=2562026.05 | 40.7 | |
| LLaVA-Next-LLaMA-3-8BKV Size=7842026.05 | 40 | |
| Cambrian-4B + TokenCompre. + Gaze attentionKV Size=36+42026.05 | 40 | |
| SPHINX-7B2026.05 | 38.7 | |
| LLaVA-VTLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 36.7 | |
| VISTAVLM=Llama3.2-Vision, Setting=base2025.12 | 35.33 | |
| Qwen2.5-VL-3B + InfiniPot-VKV Size=1282026.05 | 35.3 | |
| Cambrian-8B-DINOv2-LKV Size=5762026.05 | 34.7 | |
| LLaVA-OV-7B + HERMESKV Size=402026.05 | 34 | |
| LLaVA-1.5 (V-13B)LLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=Visual Embeddings (E)2024.03 | 33.3 | |
| LLaVA-CapLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 32 | |
| E2EVLM=Llama3.2-Vision, Setting=SFT2025.12 | 32 | |
| LACINGparameter_scale=7B2024.11 | 32 | |
| Cambrian-8B-CLIP-LKV Size=5762026.05 | 31.3 | |
| LLaVA-DCapLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 30.7 | |
| LLaVA-SGLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 30 | |
| Vicuna-VTLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 26.7 | |
| LLaVA-1.5parameter_scale=7B2024.11 | 26 | |
| VCDparameter_scale=7B2024.11 | 26 | |
| Qwen2.5-VL-3B + InfiniPot-VKV Size=502026.05 | 24.7 | |
| BLIP-2-7BKV Size=5762026.05 | 23.3 | |
| LLaVA-1.5 (V-7B)LLM=Vicuna-7B, #PT=558K, #IT=665K, Representation=Visual Embeddings (E)2024.03 | 20.7 | |
| Vicuna-DCapLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 13.3 | |
| Vicuna-CapLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 12 | |
| Vicuna-SGLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 11.3 |