Multimodal Understanding on MMMU (Efficiency Metrics)
67.8MMMU ScoreLLaVA-1.5
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLaVA-1.5Params=-2025.06 | 67.8 | — | — | — | |
| EMMAParams=4B2025.12 | 62.5 | — | — | — | |
| KVCapsuleVLM Backbone=llama3-llava-next-8b-hf2026.05 | 60 | — | — | — | |
| MetaQuery-XL#LLM=7B2025.12 | 58.6 | — | — | — | |
| STAR-7B#LLM=7B2025.12 | 58.6 | — | — | — | |
| BLIP3-oParams=7B2025.12 | 58.6 | — | — | — | |
| UniWorld-V1Params=7B2025.12 | 58.6 | — | — | — | |
| MetaQuery-XLParams=7B + 1.6B*2025.06 | 58.6 | — | — | — | |
| BLIP3-o 8BParams=7B + 1.4B*2025.06 | 58.6 | — | — | — | |
| UniWorld-V1Params=7B + 12B*2025.06 | 58.6 | — | — | — | |
| Qwen2.5-VL#Params=7B, Model Category=Vision Language Models, Co-training=false2026.05 | 58.6 | — | — | — | |
| Qwen2.5VL 7B# Vis tok.=14002025.04 | 58 | — | — | — | |
| Kimi-VL#Params=3B/16B, Model Category=Vision Language Models, Co-training=false2026.05 | 57 | — | — | — | |
| Bagel#LLM=14B2025.12 | 55.3 | — | — | — | |
| BAGELParams=7B2025.12 | 55.3 | — | — | — | |
| BAGALType=Und., # Params=14B2025.03 | 55.3 | — | — | — | |
| BAGELParams=7B + 7B*2025.06 | 55.3 | — | — | — | |
| BAGEL#Params=7B MoT, Model Category=Unified Multimodal Large Language Models, Co-training=false2026.05 | 55.3 | — | — | — | |
| Qwen2-VL#Params=7B, Model Category=Vision Language Models, Co-training=false2026.05 | 54.1 | — | — | — | |
| GPT-4V-1106Params=-2026.05 | 53.8 | — | — | — | |
| Qwen2-VLModel Size=7B, Training Strategy=Original2026.02 | 53.7 | — | — | — | |
| UAM#Params=7B MoT, Model Category=Vision-Language-Action Models, Co-training=false2026.05 | 53.7 | — | — | — | |
| STAR-3B#LLM=3B2025.12 | 53.1 | — | — | — | |
| OmniGen2Params=3B2025.12 | 53.1 | — | — | — | |
| OmniGen2Params=3B + 4B*2025.06 | 53.1 | — | — | — | |
| Qwen2.5-VL# LLM Params=3B2026.05 | 53.1 | — | — | — | |
| Qwen2.5-VL#Params=3B, Model Category=Vision Language Models, Co-training=false2026.05 | 53.1 | — | — | — | |
| OriginalCompression Ratio (CR)=None2026.02 | 52.6 | — | — | — | |
| CREMModel Size=7B, Training Strategy=Unified framework2026.02 | 52.1 | — | — | — | |
| Baseline: all KVVLM Backbone=Qwen3-VL-8B-Instruct2026.05 | 52 | — | — | — | |
| CREM_GModel Size=7B, Training Strategy=Fine-tuned on ShareGPT-4V2026.02 | 51.7 | — | — | — | |
| Qwen2.5VL 3B# Vis tok.=14002025.04 | 51.2 | — | — | — | |
| Ovis-U1#LLM=1.5B2025.12 | 51.1 | — | — | — | |
| LLaVA-NeXTType=Und., # Params=34B2025.03 | 51.1 | — | — | — | |
| LLaVA-NeXTParams=-2025.06 | 51.1 | — | — | — | |
| DeepSeek-VL2#Params=4B/27B, Model Category=Vision Language Models, Co-training=false2026.05 | 51.1 | — | — | — | |
| Qwen2.5-VL-VIFParams=7B, Base model=Qwen-2.5-VL2026.05 | 50.99 | — | — | — | |
| Qwen2-VL-7bBitwidth=fp162026.02 | 50.6 | — | — | — | |
| BLIP3-o#LLM=8B2025.12 | 50.6 | — | — | — | |
| Qwen2.5-VLParams=7B, Base model=Qwen-2.5-VL2026.05 | 50.22 | — | — | — | |
| MBQModel=Qwen2-VL-7b, Bitwidth=W8A82026.02 | 50.1 | — | — | — | |
| LLaVA-OneVisionModel Size=8B, Redundancy level (+R)=25%2026.05 | 49.9 | — | — | — | |
| TLQModel=Qwen2-VL-7b, Bitwidth=W8A82026.02 | 49.8 | — | — | — | |
| MUSE-3B# LLM Params=2B2026.05 | 49.8 | — | — | — | |
| MUSE-3BType=Unified, # Params=2B+1.6B2026.05 | 49.8 | — | — | — | |
| Qwen2.5-VL-SFTParams=7B, Base model=Qwen-2.5-VL2026.05 | 49.67 | — | — | — | |
| Molmo 7B# Vis tok.=12002025.04 | 49.1 | — | — | — | |
| PrumergeVLM Backbone=llama3-llava-next-8b-hf2026.05 | 49 | — | — | — | |
| Show-o2#LLM=7B2025.12 | 48.9 | — | — | — | |
| Show-o2Type=Uni., # Params=7B, Resolution=432px2025.03 | 48.9 | — | — | — | |
| Show-o2Type=Unified, # Params=7B2026.05 | 48.9 | — | — | — | |
| Gemini ProParams=-2026.05 | 48.9 | — | — | — | |
| mPLUG-Owl3Params=8B2026.05 | 48.9 | — | — | — | |
| LLava-OV#Params=7B, Model Category=Vision Language Models, Co-training=false2026.05 | 48.8 | — | — | — | |
| UniLIP-3B# LLM Params=2B2026.05 | 48.7 | — | — | — | |
| UniLIP-3BType=Unified, # Params=2B+1.6B2026.05 | 48.7 | — | — | — | |
| OpenUni-LType=Dual, # Params=2B+1.6B2026.05 | 48.6 | — | — | — | |
| InternVL3# LLM Params=1.8B2026.05 | 48.2 | — | — | — | |
| InternVL3Type=Und. Only, # Params=1.8B2026.05 | 48.2 | — | — | — | |
| KVCapsuleVLM Backbone=Qwen3-VL-8B-Instruct2026.05 | 48 | — | — | — | |
| PrumergeVLM Backbone=Qwen3-VL-8B-Instruct2026.05 | 48 | — | — | — | |
| LLaVA-OneVision 7B# Vis tok.=24002025.04 | 47.7 | — | — | — | |
| LaVerVisual Encoder Model=Qwen-ViT2025.12 | 47.56 | — | — | — | |
| DualToken-7BType=Uni., # Params=7B, Resolution=384px2025.03 | 47.4 | — | — | — | |
| BaselineModel=InternVL2-26B, FLOPs Ratio=02025.03 | 47.11 | 69.02 | 17,347 | 2.24 | |
| CREM_RModel Size=7B, Training Strategy=Trained on MMEB2026.02 | 47 | — | — | — | |
| TopVModel=InternVL2-26B, FLOPs Ratio=47%2025.03 | 46.98 | 68.26 | 15,736 | 2.47 | |
| FastVModel=InternVL2-26B, FLOPs Ratio=46%2025.03 | 46.91 | 165.23 | 23,012 | 1.69 | |
| BLIP3o-4BType=Und. and Gen. > 2B, # Total Params=7.1B2026.02 | 46.6 | — | — | — | |
| BLIP3-o# LLM Params=4B2026.05 | 46.6 | — | — | — | |
| LaVerVisual Encoder Model=SigLIP22025.12 | 46.33 | — | — | — | |
| FLARE-X 3B# Vis tok.=14002025.04 | 46.3 | — | — | — | |
| TLQModel=LLaVA-onevision-7b, Bitwidth=W8A82026.02 | 46.2 | — | — | — | |
| LLaVA-onevision-7bBitwidth=fp162026.02 | 46 | — | — | — | |
| Baseline: all KVVLM Backbone=llama3-llava-next-8b-hf2026.05 | 46 | — | — | — | |
| FastVVLM Backbone=llama3-llava-next-8b-hf2026.05 | 46 | — | — | — | |
| DualToken-7BType=Uni., # Params=7B, Resolution=256px2025.03 | 45.8 | — | — | — | |
| MiniCPM-Llama3-V-2.5 8B# Vis tok.=4002025.04 | 45.8 | — | — | — | |
| BaselineVisual Encoder Model=Qwen-ViT2025.12 | 45.67 | — | — | — | |
| MBQModel=LLaVA-onevision-7b, Bitwidth=W8A82026.02 | 45.6 | — | — | — | |
| FLARE-X 8B# Vis tok.=14002025.04 | 45.6 | — | — | — | |
| LaVerVisual Encoder Model=AIMv22025.12 | 45 | — | — | — | |
| SparseVLMsVLM Backbone=llama3-llava-next-8b-hf2026.05 | 45 | — | — | — | |
| BaselineVisual Encoder Model=SigLIP22025.12 | 44.78 | — | — | — | |
| COMPOTCompression Ratio (CR)=0.22026.02 | 44.7 | — | — | — | |
| FLARE-L 3B# Vis tok.=6302025.04 | 44.6 | — | — | — | |
| BaselineVisual Encoder Model=CLIP2025.12 | 44.56 | — | — | — | |
| LaVerVisual Encoder Model=CLIP2025.12 | 44.56 | — | — | — | |
| BaselineVisual Encoder Model=AIMv22025.12 | 44.56 | — | — | — | |
| ILLUME+#LLM=3B2025.12 | 44.3 | — | — | — | |
| MogaoParams=7B2025.12 | 44.2 | — | — | — | |
| MUSE-1B# LLM Params=1B2026.05 | 44.1 | — | — | — | |
| MUSE-1BType=Unified, # Params=1B+0.6B2026.05 | 44.1 | — | — | — | |
| FreeCorrectionbackbone=Lumina-DiMOO2026.02 | 44 | — | — | — | |
| Eagle 8B# Vis tok.=10242025.04 | 43.8 | — | — | — | |
| Florence-VL 8B# Vis tok.=5762025.04 | 43.7 | — | — | — | |
| InternVL2.5# LLM Params=1.8B2026.05 | 43.6 | — | — | — | |
| Lumina-DiMOO (ReMDM)2026.02 | 43.4 | — | — | — | |
| InternVL3# LLM Params=1B2026.05 | 43.4 | — | — | — | |
| Phi 3.5-Vision# Vis tok.=21002025.04 | 43.3 | — | — | — |