Multimodal Understanding on LLaVAW
76.1ScoreOPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| OPPOTraining=Oriented Pickup Preference Optimization, Base Model=LLaVA-NeXT2026.06 | 76.1 | |
| mDPOTraining=Preference Optimization, Base Model=LLaVA-NeXT2026.06 | 75.9 | |
| CHiPTraining=Preference Optimization, Base Model=LLaVA-NeXT2026.06 | 75.8 | |
| DPOTraining=Preference Optimization, Base Model=LLaVA-NeXT2026.06 | 75.5 | |
| SoM-LLaVA-1.5LLM=Vicuna-13B, Res.=336, Pre-Data=558K, IT-Data=695K2024.04 | 75.3 | |
| LLaVA-NeXT (Baseline)Training=Vanilla2026.06 | 74.7 | |
| AKI-4BModel Scale=4B, Model Access Type=Open-source2025.03 | 74.6 | |
| SPHINXLLM=LLAMA2-7B, Res.=2242024.04 | 73.5 | |
| SoM-LLaVA-1.5-TLLM=Vicuna-13B, Res.=336, Pre-Data=558K, IT-Data=695K, Tagged Images=true2024.04 | 73.3 | |
| MM1.5-3BModel Scale=3B, Model Access Type=Proprietary2025.03 | 73 | |
| ArcanaVision Encoder=VIT-L (0.3B), Language Model=Vicuna (7B), MM LoRA=trained during pretrain stage2024.10 | 72.7 | |
| MM1-3BModel Scale=3B, Model Access Type=Proprietary2025.03 | 72.1 | |
| LLaVA-1.5LLM=Vicuna-13B, Res.=336, Pre-Data=558K, IT-Data=665K2024.04 | 70.7 | |
| BLIP-3-4BModel Scale=4B, Model Access Type=Open-source2025.03 | 69.8 | |
| MiniCPM-V2-3BModel Scale=3B, Model Access Type=Open-source2025.03 | 69.2 | |
| ArcanaVision Encoder=VIT-L (0.3B), Language Model=Vicuna (7B)2024.10 | 67.3 | |
| Honeybee-C-7BModel Scale=7B, Model Access Type=Open-source2025.03 | 67.1 | |
| VILA-1.5-3BModel Scale=3B, Model Access Type=Open-source2025.03 | 65.5 | |
| Phi-3-Vision-4BModel Scale=4B, Model Access Type=Open-source2025.03 | 63.9 | |
| LLaVA-v1.5Vision Encoder=VIT-L (0.3B), Language Model=Vicuna (7B)2024.10 | 63.4 | |
| LLaVAVision Encoder=VIT-L (0.3B), Language Model=Vicuna (7B)2024.10 | 63 | |
| LLaVA-1.5-7BModel Scale=7B, Model Access Type=Open-source2025.03 | 61.8 | |
| InstructBLIPVision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B)2024.10 | 60.9 | |
| InstructBLIPLLM=Vicuna-7B, Res.=224, Pre-Data=129M, IT-Data=1.2M2024.04 | 60.9 | |
| InstructBLIPLLM=Vicuna-13B, Res.=224, Pre-Data=129M, IT-Data=1.2M2024.04 | 58.2 | |
| DeepSeek-VL-1.3BModel Scale=1.3B, Model Access Type=Open-source2025.03 | 51.1 | |
| Qwen2-VL-2BModel Scale=2B, Model Access Type=Open-source2025.03 | 50.5 | |
| MiniGPT-4Vision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B)2024.10 | 45.1 | |
| BLIP-2LLM=Vicuna-13B, Res.=224, Pre-Data=129M2024.04 | 38.1 |