Multimodal Understanding and Reasoning on MME
78.7MME ScoreLLaVA-NeXT
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaVA-NeXTLLM=Vicuna-1.5-13B, #Sample=1.3M, #Param=≥7B2026.01 | 78.7 | |
| VILA-7BLLM=LLaMA-7B, #Sample=50M, #Param=≥7B2026.01 | 76.7 | |
| LLaVA-1.5-7BLLM=Vicuna-1.5-7B, #Sample=1.2M, #Param=≥7B2026.01 | 75.5 | |
| LLaVA-FA-7BLLM=InternLM-2-20B, #Sample=5M, #Param=≥7B2026.01 | 74.5 | |
| Qwen-VL-ChatLLM=Qwen-7B, #Sample=1450M, #Param=≥7B2026.01 | 74.4 | |
| Bunny-3BLLM=Phi-2-2.7B, #Sample=2.7M, #Param=∼3B2026.01 | 74.4 | |
| Deepseek-VL-7BLLM=DLLM-7B, #Sample=2000M, #Param=≥7B2026.01 | 73.4 | |
| Imp-3BLLM=Phi-2-2.7B, #Sample=1.6M, #Param=∼3B2026.01 | 72.3 | |
| VILA-3BLLM=LLaMA-2.7B, #Sample=51M, #Param=∼3B2026.01 | 72.1 | |
| MobileVLMv2LLM=MLLaMA-2.7B, #Sample=3.6M, #Param=∼3B2026.01 | 72 | |
| CogVLMLLM=Vicuna-7B, #Sample=1500M, #Param=≥7B2026.01 | 71.8 | |
| MoE-LLaVA-3BLLM=Phi-2-2.7B, #Sample=2.2M, #Param=∼3B2026.01 | 71.1 | |
| LLaVA-FA-3BLLM=LLaMA-3-8B, #Sample=5M, #Param=∼3B2026.01 | 71 | |
| MiniCPM-V-2LLM=MiniCPM-2.4B, #Sample=570M, #Param=∼3B2026.01 | 70.5 | |
| MiniCPM-VLLM=MiniCPM-2.4B, #Sample=570M, #Param=∼3B2026.01 | 68.9 | |
| Mini-Gemini-2BLLM=Gemma-2B, #Sample=2.7M, #Param=∼2B2026.01 | 67 | |
| LLaVA-FA-2BLLM=Qwen-2.5-7B, #Sample=5M, #Param=∼2B2026.01 | 66.6 | |
| DeepSeek-VL-1.3BLLM=DLLM-1.3B, #Sample=2000M, #Param=∼2B2026.01 | 65.3 | |
| Imp-2BLLM=Qwen-1.5-1.8B, #Sample=1.6M, #Param=∼2B2026.01 | 65.2 | |
| Bunny-2BLLM=Qwen-1.5-1.8B, #Sample=2.7M, #Param=∼2B2026.01 | 65 | |
| BLIP-2LLM=Vicuna-13B, #Sample=129M, #Param=≥7B2026.01 | 64.7 | |
| MoE-LLaVA-2BLLM=Qwen-1.5-1.8B, #Sample=2.2M, #Param=∼2B2026.01 | 64.6 | |
| MobileVLMLLM=MLLaMA-2.7B, #Sample=1.3M, #Param=∼3B2026.01 | 64.4 | |
| LLaVA-FA-1BLLM=Qwen-2.5-3B, #Sample=5M, #Param=∼1B2026.01 | 63.3 | |
| SPHINX-TinyLLM=TLLaMA-1.1B, #Sample=15M, #Param=∼1B2026.01 | 63.1 | |
| InstructBLIPLLM=Vicuna-13B, #Sample=130M, #Param=≥7B2026.01 | 60.6 |