Multimodal Understanding on MME Perception
1,699.5MME-P ScoreVLM-UniDDT
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VLM-UniDDT# Params.=4B+1B2026.06 | 1,699.5 | — | — | |
| Qwen2.5 VLParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction, VLMEvalKit=true2026.03 | 1,692.5 | — | — | |
| WallarooParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction, VLMEvalKit=true2026.03 | 1,690.3 | — | — | |
| BAGELParams=7B+7B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 1,687 | — | — | |
| BAGEL# Params.=14B2026.06 | 1,687 | — | — | |
| MetaQueryParams=7B + 1.6B, Architecture category=Multi-modal Understanding Models as Enhanced Conditional Encoders2026.03 | 1,685.2 | — | — | |
| Blip3-oParams=7B + 1.4B, Architecture category=Multi-modal Understanding Models as Enhanced Conditional Encoders2026.03 | 1,682.6 | — | — | |
| Show-o2Params=7B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 1,620.5 | — | — | |
| Show-o2# Params.=7B2026.06 | 1,620.5 | — | — | |
| MogaoParams=7B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 1,592 | — | — | |
| Mogao# Params.=7B2026.06 | 1,592 | — | — | |
| LLaVA-OneVisionLLM=Qwen2-7B2025.05 | 1,580 | — | — | |
| LLaVA-OV# Params.=7B2026.06 | 1,580 | — | — | |
| ShareGPT4VLLM=Vicuna-7B2025.05 | 1,567.4 | — | — | |
| Janus-Pro-7BType=Und. and Gen., LLM Params=7B2025.01 | 1,567.1 | — | — | |
| Janus-ProParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 1,567.1 | — | — | |
| Janus-Pro# Params.=7B2026.06 | 1,567.1 | — | — | |
| TokenFlow-XL*# Params.=14B2026.06 | 1,551.1 | — | — | |
| TokenFlow-XLType=Und. and Gen., LLM Params=13B2025.01 | 1,545.9 | — | — | |
| TokenFlowParams=13B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 1,545.9 | — | — | |
| LLaVA-NeXTLLM=Vicuna-7B2025.05 | 1,519 | — | — | |
| LLaVA-v1.5Type=Und. Only, # LLM Params=7B, External pretrained diffusion model=false2024.10 | 1,510.7 | — | — | |
| LLaVA-v1.5Type=Und. Only, LLM Params=7B2025.01 | 1,510.7 | — | — | |
| LLaVA-v1.5# Params.=7B2026.06 | 1,510.7 | — | — | |
| Qwen-VL# Params.=7B2026.06 | 1,487.6 | — | — | |
| Qwen-VL-ChatType=Und. Only, # LLM Params=7B, External pretrained diffusion model=false2024.10 | 1,487.5 | — | — | |
| Qwen-VL-ChatType=Und. Only, LLM Params=7B2025.01 | 1,487.5 | — | — | |
| MorphTokensLLM=Vicuna-7B, PT Data Size=30M2025.05 | 1,477.7 | — | — | |
| Slot-MLLM (14B)LLM=Qwen2.5-14B-Instruct, PT Data Size=25M2025.05 | 1,468.3 | — | — | |
| Slot-MLLM (14B)LLM=Qwen2.5-14B-Instruct2025.05 | 1,468.3 | — | — | |
| UniTokLLM=LLaMA-2-7B, PT Data Size=70M2025.05 | 1,465.7 | — | — | |
| Liquid# Params.=8B2026.06 | 1,448 | — | — | |
| ILLUMEType=Und. and Gen., LLM Params=7B2025.01 | 1,445.3 | — | — | |
| ILLUME# Params.=7B2026.06 | 1,445.3 | — | — | |
| Janus-Pro-1BType=Und. and Gen., LLM Params=1.5B2025.01 | 1,444 | — | — | |
| Janus-Pro# Params.=1.5B2026.06 | 1,444 | — | — | |
| MobileVLM-V2Type=Und. Only, # LLM Params=2.7B, External pretrained diffusion model=false2024.10 | 1,440.5 | — | — | |
| MobileVLM-V2Type=Und. Only, LLM Params=2.7B2025.01 | 1,440.5 | — | — | |
| VILA-UType=Und. and Gen., # LLM Params=7B, External pretrained diffusion model=false2024.10 | 1,401.8 | — | — | |
| VILA-UType=Und. and Gen., LLM Params=7B2025.01 | 1,401.8 | — | — | |
| VILA-UParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 1,401.8 | — | — | |
| VILA-U# Params.=7B2026.06 | 1,401.8 | — | — | |
| JanusType=Und. and Gen., # LLM Params=1.3B, External pretrained diffusion model=false2024.10 | 1,338 | — | — | |
| JanusType=Und. and Gen., LLM Params=1.5B2025.01 | 1,338 | — | — | |
| JanusParams=1.5B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 1,338 | — | — | |
| VILA-ULLM=LLaMA-2-7B, PT Data Size=23M2025.05 | 1,336.2 | — | — | |
| LLaVA-PhiType=Und. Only, # LLM Params=2.7B, External pretrained diffusion model=false2024.10 | 1,335.1 | — | — | |
| LLaVA-PhiType=Und. Only, LLM Params=2.7B2025.01 | 1,335.1 | — | — | |
| JanusFlowParams=1.3B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 1,333.1 | — | — | |
| JanusFlow# Params.=1.5B2026.06 | 1,333.1 | — | — | |
| MobileVLM-V2Type=Und. Only, # LLM Params=1.4B, External pretrained diffusion model=false2024.10 | 1,302.8 | — | — | |
| MobileVLM-V2Type=Und. Only, LLM Params=1.4B2025.01 | 1,302.8 | — | — | |
| MobileVLMType=Und. Only, # LLM Params=2.7B, External pretrained diffusion model=false2024.10 | 1,288.9 | — | — | |
| MobileVLMType=Und. Only, LLM Params=2.7B2025.01 | 1,288.9 | — | — | |
| Emu3-ChatType=Und. Only, LLM Params=8B2025.01 | 1,244 | — | — | |
| Slot-MLLM (7B)LLM=Vicuna-7B, PT Data Size=25M2025.05 | 1,225.2 | — | — | |
| Slot-MLLM (7B)LLM=Vicuna-7B2025.05 | 1,225.2 | — | — | |
| InstructBLIPType=Und. Only, # LLM Params=13B, External pretrained diffusion model=false2024.10 | 1,212.8 | — | — | |
| InstructBLIPType=Und. Only, LLM Params=13B2025.01 | 1,212.8 | — | — | |
| MobileVLMType=Und. Only, # LLM Params=1.4B, External pretrained diffusion model=false2024.10 | 1,196.2 | — | — | |
| MobileVLMType=Und. Only, LLM Params=1.4B2025.01 | 1,196.2 | — | — | |
| LLaVA-v1.5-Phi-1.5Type=Und. Only, # LLM Params=1.3B, External pretrained diffusion model=false2024.10 | 1,128 | — | — | |
| LLaVA-v1.5-Phi-1.5Type=Und. Only, LLM Params=1.3B2025.01 | 1,128 | — | — | |
| D-DitType=Und. and Gen., LLM Params=2.0B2025.01 | 1,124.7 | — | — | |
| SEED-LLaMALLM=Vicuna-7B, PT Data Size=77M2025.05 | 1,123.9 | — | — | |
| Show-oType=Und. and Gen., LLM Params=1.3B, Resolution=5122025.01 | 1,097.2 | — | — | |
| Show-oParams=1.3B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 1,097.2 | — | — | |
| Show-o# Params.=1.3B2026.06 | 1,097.2 | — | — | |
| LaVITLLM=LLaMA-7B, PT Data Size=100M2025.05 | 964.5 | — | — | |
| Show-oType=Und. and Gen., # LLM Params=1.3B, External pretrained diffusion model=false2024.10 | 948.4 | — | — | |
| Show-oType=Und. and Gen., LLM Params=1.3B, Resolution=2562025.01 | 948.4 | — | — | |
| LLAVAType=Und. Only, # LLM Params=7B, External pretrained diffusion model=false2024.10 | 809.6 | — | — | |
| LLaVAType=Und. Only, LLM Params=7B2025.01 | 809.6 | — | — | |
| LWMLLM=LLaMA-2-7B, PT Data Size=-2025.05 | 44.8 | — | — | |
| AKI-4BModel Scale=4B, Model Access Type=Open-source2025.03 | — | 1,491.9 | — | |
| Bagel#LLM=14B2025.12 | — | — | 1,687 | |
| BLIP-2zero-shot=true2023.05 | — | 290 | — | |
| BLIP-3-4BModel Scale=4B, Model Access Type=Open-source2025.03 | — | 1,487.6 | — | |
| BLIP3-o#LLM=8B2025.12 | — | — | 1,682.6 | |
| Bunnynumber of parameters=8B2025.07 | — | — | 1,987.7 | |
| Cambrian-1Size=8B, Type=AR, Samples=-2025.12 | — | 1,547 | — | |
| ChatVLAnumber of parameters=1.5B2025.07 | — | — | 1,435.2 | |
| COINCIDE2025.02 | — | 1,495.6 | — | |
| DataTailor2025.02 | — | 1,476.1 | — | |
| DeepSeek-VL-1.3BModel Scale=1.3B, Model Access Type=Open-source2025.03 | — | 1,306.6 | — | |
| DiffusionVLSize=3B, Type=Diff., Samples=738K2025.12 | — | 1,539 | — | |
| DiffusionVLSize=7B, Type=Diff., Samples=738K2025.12 | — | 1,519 | — | |
| DimpleSize=7B, Type=Diff., Samples=1.3M2025.12 | — | 1,514 | — | |
| Eagle2number of parameters=1.5B2025.07 | — | — | 1,572.1 | |
| ECoTnumber of parameters=7B2025.07 | — | — | 0 | |
| EL2N2025.02 | — | 1,356.5 | — | |
| EMU3#LLM=8B2025.12 | — | — | 1,243.8 | |
| Full-Finetune2025.02 | — | 1,510.7 | — | |
| GraNd2025.02 | — | 1,400.5 | — | |
| Honeybee-C-7BModel Scale=7B, Model Access Type=Open-source2025.03 | — | 1,584.2 | — | |
| ICONS2025.02 | — | 1,485.7 | — | |
| ILLUME+#LLM=3B2025.12 | — | — | 1,414 | |
| InstructionGPT-42025.02 | — | 463.3 | — | |
| InstructVLA-Generalistnumber of parameters=1.5B2025.07 | — | — | 1,529.6 | |
| InstructVLA-Generalistnumber of parameters=1.5B, robot state=true2025.07 | — | — | 1,548 |