Multimodal Understanding on MME Cognition
1,293.8ScoreBLIP-2
Evaluation Results
| Method | Links | |
|---|---|---|
| BLIP-2zero-shot=true2023.05 | 1,293.8 | |
| LLaMA-Adapter V2zero-shot=true2023.05 | 972.6 | |
| LaVINzero-shot=true2023.05 | 963.6 | |
| MiniGPT-4zero-shot=true2023.05 | 866.5 | |
| DiffusionVLSize=7B, Type=Diff., Samples=738K2025.12 | 675 | |
| Qwen2.5VLSize=7B, Type=AR, Samples=>9M2025.12 | 646 | |
| Qwen2.5VLSize=3B, Type=AR, Samples=>9M2025.12 | 620 | |
| DiffusionVLSize=3B, Type=Diff., Samples=738K2025.12 | 594 | |
| LLaVAzero-shot=true2023.05 | 502.8 | |
| LLaDA-VSize=8B, Type=Diff., Samples=16.5M2025.12 | 491 | |
| DimpleSize=7B, Type=Diff., Samples=1.3M2025.12 | 432 | |
| LyricsLLM Backbone=Vicuna-13B2023.12 | 431.6 | |
| LLaVA-OVSize=7B, Type=AR, Samples=7.8M2025.12 | 418 | |
| Qwen2-VL-2BModel Scale=2B, Model Access Type=Open-source2025.03 | 413.9 | |
| MiniCPM-V2-3BModel Scale=3B, Model Access Type=Open-source2025.03 | 396.8 | |
| ShareGPT4VLLM Backbone=Vicuna-7B2023.12 | 376.4 | |
| AKI-4BModel Scale=4B, Model Access Type=Open-source2025.03 | 362.9 | |
| Qwen-VL-7B-ChatModel Size=7B, Mode=Chat2023.11 | 360.71 | |
| Qwen-VL-ChatLLM Backbone=Qwen-7B2023.12 | 360.7 | |
| LLaMA-Adapter V22023.11 | 356.43 | |
| LaViDa-LSize=8B, Type=Diff., Samples=1.6M2025.12 | 341 | |
| PRISM2025.02 | 330 | |
| SPHINX-2kResolution=2k2023.11 | 326.8 | |
| SPHINX2023.11 | 322.2 | |
| TIVE2025.02 | 322.1 | |
| MM1.5-3BModel Scale=3B, Model Access Type=Proprietary2025.03 | 319.6 | |
| DataTailor2025.02 | 319.2 | |
| Full-Finetune2025.02 | 311.9 | |
| SPHINX-1kResolution=1k2023.11 | 310 | |
| Honeybee-C-7BModel Scale=7B, Model Access Type=Open-source2025.03 | 307.1 | |
| Length2025.02 | 306 | |
| Phi-3-Vision-4BModel Scale=4B, Model Access Type=Open-source2025.03 | 302.9 | |
| BLIP-3-4BModel Scale=4B, Model Access Type=Open-source2025.03 | 302.9 | |
| LLaVA-1.5-7BModel Scale=7B, Model Access Type=Open-source2025.03 | 302.1 | |
| LLaVA-1.5LLM Backbone=Vicuna-13B2023.12 | 295.4 | |
| LLaVA1.5-13BModel Size=13B2023.11 | 295.36 | |
| EL2N2025.02 | 294.7 | |
| BLIP-2LLM Backbone=FLAN-T52023.12 | 290 | |
| GraNd2025.02 | 287.1 | |
| MM1-3BModel Scale=3B, Model Access Type=Proprietary2025.03 | 279.3 | |
| VILA-1.5-3BModel Scale=3B, Model Access Type=Open-source2025.03 | 268.2 | |
| Perplexity2025.02 | 260.7 | |
| LLaVALLM Backbone=Vicuna-7B2023.12 | 247.9 | |
| Random2025.02 | 233.5 | |
| DeepSeek-VL-1.3BModel Scale=1.3B, Model Access Type=Open-source2025.03 | 225 |