Multi-modal Vision-Language Understanding on GQA
64.2AccuracyLLaVA-NeXT
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LLaVA-NeXTLLM=Vicuna-7B2025.05 | 64.2 | — | — | |
| TUNAType=Native Unified, # Params.=7B2026.05 | 63.9 | — | — | |
| ASPOBase Model=LLaVA-v1.5, Model Scale=13B, Optimization Protocol=ASPO2025.05 | 63.4 | 68.8 | — | |
| LLaVA-1.5-13BBase Model=LLaVA-v1.5, Model Scale=13B, Optimization Protocol=Base2025.05 | 63.3 | 65.93 | — | |
| Show-o2Params=7B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 63.1 | — | — | |
| Show-o2Type=Native Unified, # Params.=7B2026.05 | 63.1 | — | — | |
| LLaVA-NeXT 7BToken Budget=2880, Base Model=LLaVA-NeXT 7B2026.03 | 62.8 | — | 100 | |
| TokenFlowParams=13B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 62.7 | — | — | |
| DPOBase Model=LLaVA-v1.5, Model Scale=13B, Optimization Protocol=DPO2025.05 | 62.3 | 67.03 | — | |
| LLaVA-1.5-7BBase Model=LLaVA-v1.5, Model Scale=7B, Optimization Protocol=Base2025.05 | 62 | 62.59 | — | |
| ASPOBase Model=LLaVA-v1.5, Model Scale=7B, Optimization Protocol=ASPO2025.05 | 62 | 65.16 | — | |
| Janus-ProParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 62 | — | — | |
| LLaVA-v1.5Type=Und. Only, # Params.=7B2026.05 | 62 | — | — | |
| Janus-ProType=Native Unified, # Params.=7B2026.05 | 62 | — | — | |
| PromPruneToken Budget=640, Base Model=LLaVA-NeXT 7B2026.03 | 61.8 | — | 99.6 | |
| DivPruneToken Budget=640, Base Model=LLaVA-NeXT 7B2026.03 | 61.7 | — | 99.6 | |
| TUNAType=Native Unified, # Params.=1.5B2026.05 | 61.4 | — | — | |
| VisPrunerToken Budget=640, Base Model=LLaVA-NeXT 7B2026.03 | 61.3 | — | 99.2 | |
| DPOBase Model=LLaVA-v1.5, Model Scale=7B, Optimization Protocol=DPO2025.05 | 61 | 63.12 | — | |
| DivPruneToken Budget=320, Base Model=LLaVA-NeXT 7B2026.03 | 61 | — | 97.7 | |
| MogaoParams=7B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 60.9 | — | — | |
| MogaoType=Native Unified, # Params.=7B2026.05 | 60.9 | — | — | |
| VILA-UParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 60.8 | — | — | |
| PromPruneToken Budget=320, Base Model=LLaVA-NeXT 7B2026.03 | 60.7 | — | 97.9 | |
| Qwen-2.5-VL-InstructType=Und. Only, # Params.=7B2026.05 | 60.7 | — | — | |
| JanusFlowParams=1.3B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 60.3 | — | — | |
| Emu3Params=8B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 60.3 | — | — | |
| Qwen2.5 VLParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction, VLMEvalKit=true2026.03 | 60.3 | — | — | |
| Emu3Type=Native Unified, # Params.=8B2026.05 | 60.3 | — | — | |
| WallarooParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction, VLMEvalKit=true2026.03 | 60.1 | — | — | |
| DivPruneToken Budget=160, Base Model=LLaVA-NeXT 7B2026.03 | 59.5 | — | 95 | |
| Qwen-VL-7BBase Model=Qwen-VL, Model Scale=7B, Optimization Protocol=Base2025.05 | 59.3 | — | — | |
| JanusParams=1.5B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 59.1 | — | — | |
| PromPruneToken Budget=160, Base Model=LLaVA-NeXT 7B2026.03 | 59.1 | — | 95.2 | |
| Slot-MLLM (14B)LLM=Qwen2.5-14B-Instruct2025.05 | 58.8 | — | — | |
| VisPrunerToken Budget=320, Base Model=LLaVA-NeXT 7B2026.03 | 58.7 | — | 95.9 | |
| Show-oParams=1.3B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 58 | — | — | |
| Qwen-VL-chat-7BBase Model=Qwen-VL-Chat, Model Scale=7B, Optimization Protocol=Base2025.05 | 57.5 | — | — | |
| Qwen-VL-ChatType=Und. Only, # Params.=7B2026.05 | 57.5 | — | — | |
| Slot-MLLM (7B)LLM=Vicuna-7B2025.05 | 57.3 | — | — | |
| mPLUG-Owl2-7BBase Model=mPLUG-Owl2, Model Scale=7B, Optimization Protocol=Base2025.05 | 56.1 | — | — | |
| STARFlow2Type=Native Unified, # Params.=10.6B2026.05 | 55.8 | — | — | |
| VisPrunerToken Budget=160, Base Model=LLaVA-NeXT 7B2026.03 | 55.6 | — | 90.6 | |
| InstructBLIPLLM=Vicuna-7B2025.05 | 49.2 | — | — | |
| SEED-XType=Composite Unified, # Params.=17B2026.05 | 49.1 | — | — | |
| ASPOBase Model=InstructBLIP, Model Scale=13B, Optimization Protocol=ASPO2025.05 | 48.2 | 45.84 | — | |
| InstructBLIP-13BBase Model=InstructBLIP, Model Scale=13B, Optimization Protocol=Base2025.05 | 48.1 | 43.86 | — | |
| DPOBase Model=InstructBLIP, Model Scale=13B, Optimization Protocol=DPO2025.05 | 47.5 | 44.5 | — | |
| IDEFICS-65BBase Model=IDEFICS, Model Scale=65B, Optimization Protocol=Base2025.05 | 45.2 | — | — | |
| BLIP-2-7BBase Model=BLIP-2, Model Scale=7B, Optimization Protocol=Base2025.05 | 41 | — | — | |
| IDEFICS-7BBase Model=IDEFICS, Model Scale=7B, Optimization Protocol=Base2025.05 | 38.4 | — | — |