Multi-modal Understanding on LLaVA-Bench Wild
91.2LLaVA^W ScoreGPT4V
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT4VPrompting Strategy=Compositional Chain-of-Thought (CCoT)2023.11 | 91.2 | — | — | |
| GPT4VPrompting Strategy=Zero-Shot CoT2023.11 | 88.8 | — | — | |
| Qwen-VL-Max + AutoVBase Model=Qwen-VL-Max, AutoV Integration=true, Retrieval Strategy Source=AutoV pre-trained on LLaVA-OneVision2025.06 | 88.4 | — | — | |
| GPT4VPrompting Strategy=Base2023.11 | 88.2 | — | — | |
| VILA2-8BLLM Parameters=8B, Vision Tower Parameters=400M, Tokens per image=196, Pre-training data size=51M2024.07 | 86.6 | — | — | |
| Qwen-VL-MaxBase Model=Qwen-VL-Max, AutoV Integration=false2025.06 | 85.8 | — | — | |
| ASPOBase Model=LLaVA-v1.5, Model Scale=13B, Optimization Protocol=ASPO2025.05 | 82 | 68.8 | — | |
| MM1-7B-ChatLLM Parameters=7B, Vision Tower Parameters=300M, Tokens per image=720, Pre-training data size=>2B2024.07 | 81.5 | — | — | |
| Gemini-1.5-Pro + AutoVBase Model=Gemini-1.5-Pro, AutoV Integration=true, Retrieval Strategy Source=AutoV pre-trained on LLaVA-OneVision2025.06 | 81.1 | — | — | |
| Gemini-1.5-ProBase Model=Gemini-1.5-Pro, AutoV Integration=false2025.06 | 77.9 | — | — | |
| Honeybeeprojector=C-Abstractor2023.12 | 77.5 | — | — | |
| DOP-OBCBase Model=LLaVA-1.52026.04 | 77.3 | — | — | |
| DOP-OBCBase Model=Video-LLaVA2026.04 | 76.4 | — | — | |
| DOP-OBCBase Model=Chat-UniVi2026.04 | 75.9 | — | — | |
| HoneybeeLLM=Vicuna-13B, Projector=C-Abstractor, Vision Encoder=CLIP ViT-L/14, Res.=336, Visual tokens (M)=2562023.12 | 75.7 | — | — | |
| ASPOBase Model=LLaVA-v1.5, Model Scale=7B, Optimization Protocol=ASPO2025.05 | 75.7 | 65.16 | — | |
| AlignGPTLLM=Vicuna-13B, Resolution=3362024.05 | 75.2 | — | — | |
| CAMDBase Model=LLaVA-1.52026.03 | 75 | — | — | |
| CAMDBase Model=Video-LLaVA2026.03 | 75 | — | — | |
| LLaVA-1.5-13BPrompting Strategy=Compositional Chain-of-Thought (CCoT)2023.11 | 74.9 | — | — | |
| DPOBase Model=LLaVA-v1.5, Model Scale=13B, Optimization Protocol=DPO2025.05 | 74.7 | 67.03 | — | |
| FarSightBase Model=LLaVA-1.52026.03 | 74.7 | — | — | |
| FarSightBase Model=LLaVA-1.52026.04 | 74.7 | — | — | |
| FarSightBase Model=Video-LLaVA2026.03 | 74.5 | — | — | |
| CAMDBase Model=Chat-UniVi2026.03 | 73.6 | — | — | |
| LLaVA-1.5-13BPrompting Strategy=Base2023.11 | 73.5 | — | — | |
| Video-LLaVADecoding Strategy=Greedy2026.03 | 73.1 | — | — | |
| BaseBase Model=Video-LLaVA2026.04 | 73.1 | — | — | |
| HoneybeeLLM=Vicuna-13B, Projector=D-Abstractor, Vision Encoder=CLIP ViT-L/14, Res.=336, Visual tokens (M)=2562023.12 | 72.9 | — | — | |
| FarSightBase Model=Chat-UniVi2026.03 | 72.6 | — | — | |
| LLaVA-1.5Decoding Strategy=Greedy2026.03 | 72.5 | — | — | |
| BaseBase Model=LLaVA-1.52026.04 | 72.5 | — | — | |
| OPERABase Model=LLaVA-1.52026.03 | 72 | — | — | |
| OPERABase Model=LLaVA-1.52026.04 | 72 | — | — | |
| VisionZipBase Model=LLaVA 1.5 7B, Number of Visual Tokens=1922024.12 | 71.3 | — | 106.7 | |
| CGDBase Model=LLaVA-1.52026.03 | 71.3 | — | — | |
| SphinxPrompting Strategy=Compositional Chain-of-Thought (CCoT)2023.11 | 71 | — | — | |
| GPT-4o + AutoVBase Model=GPT-4o, AutoV Integration=true, Retrieval Strategy Source=AutoV pre-trained on LLaVA-OneVision2025.06 | 70.9 | — | — | |
| VCDBase Model=LLaVA-1.52026.03 | 70.9 | — | — | |
| VCDBase Model=LLaVA-1.52026.04 | 70.9 | — | — | |
| LLaVA-1.5LLM=Vicuna-13B, Projector=Linear, Vision Encoder=CLIP ViT-L/14, Res.=3362023.12 | 70.7 | — | — | |
| LLaVA-v1.52023.12 | 70.7 | — | — | |
| LLaVA-1.5-13BLLM=Vicuna-1.5-13B, Resolution=3362024.06 | 70.7 | — | — | |
| LLaVA-1.5LLM=Vicuna-13B, Resolution=3362024.05 | 70.7 | — | — | |
| LLaVA-1.5-13BBase Model=LLaVA-v1.5, Model Scale=13B, Optimization Protocol=Base2025.05 | 70.7 | 65.93 | — | |
| Chat-UniViDecoding Strategy=Greedy2026.03 | 70.4 | — | — | |
| BaseBase Model=Chat-UniVi2026.04 | 70.4 | — | — | |
| LLaVA-Instruct-13BLLM=Vicuna-1.5-13B, Resolution=3362024.06 | 70.1 | — | — | |
| SphinxPrompting Strategy=Base2023.11 | 70 | — | — | |
| SphinxPrompting Strategy=Zero-Shot CoT2023.11 | 69.8 | — | — | |
| ICDBase Model=LLaVA-1.52026.03 | 69.7 | — | — | |
| ICDBase Model=LLaVA-1.52026.04 | 69.7 | — | — | |
| LLaVA-1.5-13BPrompting Strategy=Zero-Shot CoT2023.11 | 68.5 | — | — | |
| AlignGPTLLM=Vicuna-7B, Resolution=3362024.05 | 68.4 | — | — | |
| GPT-4oBase Model=GPT-4o, AutoV Integration=false2025.06 | 68 | — | — | |
| ASPOBase Model=InstructBLIP, Model Scale=13B, Optimization Protocol=ASPO2025.05 | 67.4 | 45.84 | — | |
| HoneybeeLLM=Vicuna-7B, Projector=C-Abstractor, Vision Encoder=CLIP ViT-L/14, Res.=224, Visual tokens (M)=1442023.12 | 67.1 | — | — | |
| VanillaBase Model=LLaVA 1.5 7B, Number of Visual Tokens=5762024.12 | 66.8 | — | 100 | |
| VisionZipBase Model=LLaVA 1.5 7B, Number of Visual Tokens=1282024.12 | 66.7 | — | 99.9 | |
| HoneybeeLLM=Vicuna-7B, Projector=D-Abstractor, Vision Encoder=CLIP ViT-L/14, Res.=224, Visual tokens (M)=1442023.12 | 66.3 | — | — | |
| DPOBase Model=LLaVA-v1.5, Model Scale=7B, Optimization Protocol=DPO2025.05 | 65.7 | 63.12 | — | |
| LLaVA-Instruct-7BLLM=Vicuna-1.5-7B, Resolution=3362024.06 | 65.1 | — | — | |
| DOP-OBCBase Model=InstructBLIP2026.04 | 64.3 | — | — | |
| VisionZipBase Model=LLaVA 1.5 7B, Number of Visual Tokens=642024.12 | 63.5 | — | 95.1 | |
| LLaVA-1.5LLM=Vicuna-7B, Projector=Linear, Vision Encoder=CLIP ViT-L/14, Res.=3362023.12 | 63.4 | — | — | |
| LLaVA-1.5-7BLLM=Vicuna-1.5-7B, Resolution=3362024.06 | 63.4 | — | — | |
| LLaVA-1.5LLM=Vicuna-7B, Resolution=3362024.05 | 63.4 | — | — | |
| LLaVA-1.5-7BBase Model=LLaVA-v1.5, Model Scale=7B, Optimization Protocol=Base2025.05 | 63.4 | 62.59 | — | |
| LLaVA-7BBase Model=LLaVA, Model Scale=7B, Optimization Protocol=Base2025.05 | 63 | — | — | |
| DPOBase Model=InstructBLIP, Model Scale=13B, Optimization Protocol=DPO2025.05 | 62.7 | 44.5 | — | |
| CAMDBase Model=InstructBLIP2026.03 | 61.2 | — | — | |
| FarSightBase Model=InstructBLIP2026.03 | 61 | — | — | |
| InstructBLIPLLM=Vicuna-7B, Projector=Q-former, Vision Encoder=EVA-CLIP ViT-G, Res.=2242023.12 | 60.9 | — | — | |
| InstructBLIP-7BLLM=Vicuna-7B, Resolution=2242024.06 | 60.9 | — | — | |
| InstructBLIPLLM=Vicuna-7B, Resolution=2242024.05 | 60.9 | — | — | |
| InstructBLIPLLM=Vicuna-13B, Projector=Q-former, Vision Encoder=EVA-CLIP ViT-G, Res.=2242023.12 | 58.2 | — | — | |
| InstructBLIPLLM=Vicuna-13B, Resolution=2242024.05 | 58.2 | — | — | |
| InstructBLIP-13BBase Model=InstructBLIP, Model Scale=13B, Optimization Protocol=Base2025.05 | 58.2 | 43.86 | — | |
| InstructBLIPDecoding Strategy=Greedy2026.03 | 58.2 | — | — | |
| BaseBase Model=InstructBLIP2026.04 | 58.2 | — | — | |
| InstructBLIP-13BPrompting Strategy=Compositional Chain-of-Thought (CCoT)2023.11 | 47.9 | — | — | |
| InstructBLIP-13BPrompting Strategy=Base2023.11 | 47.2 | — | — | |
| InstructBLIP-13BPrompting Strategy=Zero-Shot CoT2023.11 | 45.4 | — | — | |
| BLIP-2LLM=Vicuna-13B, Projector=Q-former, Vision Encoder=EVA-CLIP ViT-G, Res.=2242023.12 | 38.1 | — | — | |
| BLIP-2LLM=Vicuna-13B, Resolution=2242024.05 | 38.1 | — | — | |
| BLIP-2-7BBase Model=BLIP-2, Model Scale=7B, Optimization Protocol=Base2025.05 | 38.1 | — | — |