Visual Understanding on MME
2,321MME ScoreQwen2-VL
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2-VLModel Size=8B2025.01 | 2,321 | |
| InternVL2.5-MPOModel Size=8B2025.01 | 2,321 | |
| SAIL-VLModel Size=8B2025.01 | 2,244 | |
| DeepSeekVL-2Model Size=8B2025.01 | 2,149 | |
| InternVL2.5-MPOModel Size=2B2025.01 | 2,123 | |
| SAIL-VLModel Size=2B2025.01 | 1,969 | |
| ShareGPT4VLLM=Vicuna-7B, Resolution=3362026.03 | 1,943.8 | |
| Qwen2-VLModel Size=2B2025.01 | 1,923 | |
| DeepSeekVL-2Model Size=2B2025.01 | 1,910 | |
| EvoTokLLM=Qwen-2.5-7B-Inst., Resolution=2562026.03 | 1,895.1 | |
| InternVL3.5-1B# Total Params=1.1B2025.11 | 1,894 | |
| Qwen-VL-ChatLLM=Qwen-7B, Resolution=4482026.03 | 1,848.3 | |
| LLaVA-v1.5LLM=Vicuna-1.5-13B, Resolution=3362026.03 | 1,826.7 | |
| LFM2-VL-1.6B# Total Params=1.6B2025.11 | 1,757 | |
| Emu2-ChatType=Understanding-only, LLM Params=33B2025.03 | 1,678 | |
| TokenFlow-BLLM=Vicuna-13B, Resolution=2242026.03 | 1,660.4 | |
| TokenFlow-LLLM=Vicuna-13B, Resolution=2562026.03 | 1,622.9 | |
| Dynamic-LLaVA-13B^ITFree=false, Image (prefill) Token=115, TFLOPs=4.72024.12 | 1,563.3 | |
| Dynamic-LLaVA-13B^VFree=false, Image (prefill) Token=115, TFLOPs=4.72024.12 | 1,554.1 | |
| VILALLM=LLaMA-2-7B, Resolution=3362026.03 | 1,533 | |
| LLaVA-1.5-13BImage (prefill) Token=576, TFLOPs=19.62024.12 | 1,531.3 | |
| LLaVA-1.5-7BImage (prefill) Token=576, TFLOPs=10.12024.12 | 1,510.7 | |
| DualTokenLLM=LLaMA-2-7B, Resolution=2562026.03 | 1,502.7 | |
| Dynamic-LLaVA-7B^ITFree=false, Image (prefill) Token=115, TFLOPs=2.52024.12 | 1,501 | |
| QLIPLLM=Vicuna-1.5-7B, Resolution=3922026.03 | 1,498.3 | |
| LLaVA-PruMerge+ (13B)Image (prefill) Token=146, TFLOPs=4.92024.12 | 1,485.5 | |
| Dynamic-LLaVA-7B^VFree=false, Image (prefill) Token=115, TFLOPs=2.52024.12 | 1,479.8 | |
| LLaVA-FastV (13B)Image (prefill) Token=144, TFLOPs=62024.12 | 1,470.3 | |
| LLaVA-PruMerge+ (7B)Image (prefill) Token=146, TFLOPs=2.52024.12 | 1,462.4 | |
| LLaVA-FastV (7B)Free=true, Image (prefill) Token=144, TFLOPs=3.22024.12 | 1,458.9 | |
| LLaVA-1.5-7B+H2Ok-10, -0.5Image (prefill) Token=576, TFLOPs=10.12024.12 | 1,458.4 | |
| SmolVLM2-500M# Total Params=0.5B2025.11 | 1,456 | |
| LLaVA-1.5-13b+H2Ok-10,-0.5Image (prefill) Token=576, TFLOPs=19.62024.12 | 1,448.3 | |
| UniTokLLM=LLaMA-2-7B, Resolution=2562026.03 | 1,448 | |
| VILA-ULLM=LLaMA-7B, Resolution=3842026.03 | 1,401.8 | |
| JanusLLM=DeepSeek-1.3B, Resolution=3842026.03 | 1,338 | |
| VILA-ULLM=LLaMA-7B, Resolution=2562026.03 | 1,336.2 | |
| LLaVA-FastV† (7B)Free=false, Image (prefill) Token=144, TFLOPs=3.22024.12 | 1,292.2 | |
| LFM2-VL-450M# Total Params=0.45B2025.11 | 1,230 | |
| SeGroSFine-Tuning=Ours2026.03 | 1,217 | |
| SFTFine-Tuning=SFT2026.03 | 1,210 | |
| Base ModelFine-Tuning=w/o SFT2026.03 | 1,195 | |
| RecaFine-Tuning=Reca2026.03 | 1,190 | |
| IdeficsType=Understanding-only, LLM Params=9B2025.03 | 1,177.3 | |
| MiniGPT4Type=Understanding-only, LLM Params=7B2025.03 | 1,047.4 | |
| Kosmos-2Type=Understanding-only, LLM Params=2B2025.03 | 721.1 | |
| LLaVA-1.5-7B+H2O-0.5Image (prefill) Token=576, TFLOPs=11.62024.12 | 500.5 | |
| LaViDa-O2025.12 | 488 | |
| Qwen-VLType=Understanding-only, LLM Params=7B2025.03 | 482.7 | |
| Sparse-LaViDa2025.12 | 450 | |
| WeGenType=Unified (Understanding & Generation), LLM Params=7B2025.03 | 447.4 | |
| Chameleon-7bType=Unified (Understanding & Generation), LLM Params=7B2025.03 | 202.7 | |
| Kosmos-GType=Unified (Understanding & Generation), LLM Params=1.9B2025.03 | 104.3 | |
| LLaVAType=Understanding-only, LLM Params=7B2025.03 | 28.3 |