Multimodal Understanding on MMMU (MMMU Score)
60.74MMMU ScoreHeadLens
Evaluation Results
| Method | Links | |
|---|---|---|
| HeadLensBackbone=Qwen3-VL-8B2026.03 | 60.74 | |
| MetaQueryParams=7B + 1.6B, Architecture category=Multi-modal Understanding Models as Enhanced Conditional Encoders2026.03 | 58.6 | |
| UniWorld-V1Params=7B + 12B, Architecture category=Multi-modal Understanding Models as Enhanced Conditional Encoders2026.03 | 58.6 | |
| BaselineBackbone=Qwen3-VL-8B2026.03 | 58.33 | |
| AD-Loop#Params=7B, Capabilities=Und. and Gen.2026.02 | 57.3 | |
| HeadLensBackbone=Qwen2.5-VL-7B2026.03 | 56.33 | |
| BAGELParams=7B+7B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 55.3 | |
| BAGEL#Params=7B, Capabilities=Und. and Gen.2026.02 | 55.3 | |
| BaselineBackbone=Qwen2.5-VL-7B2026.03 | 54.89 | |
| OmniGen2Params=3B + 4B, Architecture category=Multi-modal Understanding Models as Enhanced Conditional Encoders2026.03 | 53.1 | |
| Ming-Lite-UniParams=8B+1.6B, Architecture category=Multi-modal Understanding Models as Enhanced Conditional Encoders2026.03 | 51.2 | |
| Blip3-oParams=7B + 1.4B, Architecture category=Multi-modal Understanding Models as Enhanced Conditional Encoders2026.03 | 50.6 | |
| Intern-VL-2.5-4BModel Category=Autoregressive2026.04 | 50 | |
| Show-o2Params=7B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 48.9 | |
| LLaDA-VModel Category=Diffusion2026.04 | 48.6 | |
| Qwen2.5-VL-3BModel Category=Autoregressive2026.04 | 47.3 | |
| Fast-dVLMModel Category=Diffusion, Decoding Strategy=spec.2026.04 | 46.6 | |
| DimpleModel Category=Diffusion2026.04 | 45.2 | |
| Qwen2.5 VLParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction, VLMEvalKit=true2026.03 | 44.9 | |
| Fast-dVLMModel Category=Diffusion, Decoding Strategy=MDM2026.04 | 44.6 | |
| HeadLensBackbone=LLaVA-1.5-7B2026.03 | 44.56 | |
| BaselineBackbone=LLaVA-1.5-7B2026.03 | 44.44 | |
| MogaoParams=7B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 44.2 | |
| Bunny-v1.1-Llama3Parameters=8B2026.03 | 43.3 | |
| LaViDaModel Category=Diffusion2026.04 | 43.3 | |
| TokLIP# Params=7B, Resolution=3842025.03 | 43.1 | |
| WallarooParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction, VLMEvalKit=true2026.03 | 42.7 | |
| Qwen2-VLParameters=2B2026.03 | 41.1 | |
| SemHiTok# Params=7B, Resolution=3842025.03 | 41 | |
| Janus-ProParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 41 | |
| Janus-Pro#Params=7B, Capabilities=Und. and Gen.2026.02 | 41 | |
| SemHiTok# Params=7B, Resolution=2562025.03 | 39.3 | |
| Janus-Pro-1B# Params=1.5B, Resolution=3842025.03 | 38.9 | |
| MAR# Params=1.5B, Resolution=5122025.03 | 38.9 | |
| TokenFlowParams=13B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 38.7 | |
| MiniCPM-V 2.0Parameters=2B2026.03 | 38.2 | |
| MiniCPM-V-2 (3B)Model Category=Autoregressive2026.04 | 37.9 | |
| ShareGPT4V# Params=7B, Resolution=3362025.03 | 37.2 | |
| VisionPanguParameters=1.7B2026.03 | 36.5 | |
| InternVL2Parameters=2B2026.03 | 36.3 | |
| LLaVA-v1.6-VicunaParameters=7B2026.03 | 35.8 | |
| GPrune-LLM (Wanda-sp)Ratio=90%2026.03 | 35.67 | |
| LLaVA-v1.5# Params=7B, Resolution=3362025.03 | 35.4 | |
| LLaVA-v1.5#Params=7B, Capabilities=Und. Only2026.02 | 35.4 | |
| LLaVA1.5-7BRatio=100%2026.03 | 35.11 | |
| TokenFlow-L# Params=13B, Resolution=2562025.03 | 34.4 | |
| TokenFlow-B# Params=13B, Resolution=2242025.03 | 34.2 | |
| SynerGen-VL# Params=2.4B, Resolution=5122025.03 | 34.2 | |
| GPrune-LLM (FLAP)Ratio=90%2026.03 | 33.78 | |
| UniToken# Params=7B, Resolution=3842025.03 | 32.8 | |
| FLAPRatio=90%2026.03 | 32.78 | |
| VILA-1.5-3BModel Category=Autoregressive2026.04 | 31.8 | |
| EMU3# Params=8B, Resolution=5122025.03 | 31.6 | |
| Emu3Params=8B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 31.6 | |
| Emu3#Params=8B, Capabilities=Und. and Gen.2026.02 | 31.6 | |
| Qwen-Vl-Chat# Params=7B, Resolution=4482025.03 | 30.5 | |
| Janus# Params=1.5B, Resolution=3842025.03 | 30.5 | |
| JanusParams=1.5B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 30.5 | |
| GPrune-LLM (FLAP)Ratio=80%2026.03 | 30.44 | |
| MMaDA#Params=8B, Capabilities=Und. and Gen.2026.02 | 30.2 | |
| Wanda-spRatio=90%2026.03 | 29.89 | |
| GPrune-LLM (Wanda-sp)Ratio=80%2026.03 | 29.56 | |
| JanusFlowParams=1.3B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 29.3 | |
| FLAPRatio=80%2026.03 | 28.89 | |
| Wanda-spRatio=80%2026.03 | 28.56 | |
| Show-oParams=1.3B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 27.4 | |
| Show-o# Params=1.5B, Resolution=2562025.03 | 26.7 | |
| Show-o#Params=1.3B, Capabilities=Und. and Gen.2026.02 | 26.7 | |
| ChameleonParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 22.4 |