Multimodal Understanding and Reasoning on MMBench Chinese (dev)
72.8AccuracyDeepseek-VL-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Deepseek-VL-7BLLM=DLLM-7B, #Sample=2000M, #Param=≥7B2026.01 | 72.8 | |
| LLaVA-FA-7BLLM=InternLM-2-20B, #Sample=5M, #Param=≥7B2026.01 | 69.5 | |
| LLaVA-FA-3BLLM=LLaMA-3-8B, #Sample=5M, #Param=∼3B2026.01 | 68 | |
| MiniCPM-V-2LLM=MiniCPM-2.4B, #Sample=570M, #Param=∼3B2026.01 | 67.2 | |
| LLaVA-NeXTLLM=Vicuna-1.5-13B, #Sample=1.3M, #Param=≥7B2026.01 | 64.4 | |
| MiniCPM-VLLM=MiniCPM-2.4B, #Sample=570M, #Param=∼3B2026.01 | 62.7 | |
| VILA-7BLLM=LLaMA-7B, #Sample=50M, #Param=≥7B2026.01 | 61.7 | |
| LLaVA-FA-2BLLM=Qwen-2.5-7B, #Sample=5M, #Param=∼2B2026.01 | 61.7 | |
| Imp-2BLLM=Qwen-1.5-1.8B, #Sample=1.6M, #Param=∼2B2026.01 | 61.2 | |
| DeepSeek-VL-1.3BLLM=DLLM-1.3B, #Sample=2000M, #Param=∼2B2026.01 | 61 | |
| Bunny-2BLLM=Qwen-1.5-1.8B, #Sample=2.7M, #Param=∼2B2026.01 | 58.5 | |
| LLaVA-1.5-7BLLM=Vicuna-1.5-7B, #Sample=1.2M, #Param=≥7B2026.01 | 58.3 | |
| MoE-LLaVA-2BLLM=Qwen-1.5-1.8B, #Sample=2.2M, #Param=∼2B2026.01 | 57.3 | |
| Qwen-VL-ChatLLM=Qwen-7B, #Sample=1450M, #Param=≥7B2026.01 | 56.7 | |
| CogVLMLLM=Vicuna-7B, #Sample=1500M, #Param=≥7B2026.01 | 53.8 | |
| VILA-3BLLM=LLaMA-2.7B, #Sample=51M, #Param=∼3B2026.01 | 52.7 | |
| Mini-Gemini-2BLLM=Gemma-2B, #Sample=2.7M, #Param=∼2B2026.01 | 51.3 | |
| LLaVA-FA-1BLLM=Qwen-2.5-3B, #Sample=5M, #Param=∼1B2026.01 | 49.4 | |
| Imp-3BLLM=Phi-2-2.7B, #Sample=1.6M, #Param=∼3B2026.01 | 46.7 | |
| MoE-LLaVA-3BLLM=Phi-2-2.7B, #Sample=2.2M, #Param=∼3B2026.01 | 41.8 | |
| SPHINX-TinyLLM=TLLaMA-1.1B, #Sample=15M, #Param=∼1B2026.01 | 37.8 | |
| Bunny-3BLLM=Phi-2-2.7B, #Sample=2.7M, #Param=∼3B2026.01 | 37.2 |