Visual Understanding on MM-Vet
76.9MM-Vet ScoreGPT-4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4o2025.06 | 76.9 | |
| LVRPOType=Unified (>1.5B), # LLM Params=7B2026.03 | 69.5 | |
| GPT-4o-05132024.09 | 69.1 | |
| BAGEL2025.06 | 67.2 | |
| BAGELType=Unified, LLM Parameters=7B MOT2025.05 | 67.2 | |
| BAGELType=Unified (>1.5B), # LLM Params=7B2026.03 | 67.2 | |
| BAGELModel Category=Unified Understanding & Generation Model2025.10 | 67.2 | |
| UniWorld-V12025.06 | 67.1 | |
| Qwen2.5-VLType=Und. Only, LLM Parameters=7B2025.05 | 67.1 | |
| Qwen2.5-VLType=Und. Only, # LLM Params=7B2026.03 | 67.1 | |
| Kimi-VLType=Und. Only, LLM Parameters=2.8B/16B2025.05 | 66.7 | |
| Kimi-VLType=Und. Only, # LLM Params=2.8B/16B2026.03 | 66.7 | |
| BLIP3-o2025.06 | 66.6 | |
| MetaQuery2025.06 | 66.6 | |
| MetaQuery-XLType=Unified, LLM Parameters=7B2025.05 | 66.6 | |
| MetaQuery-XLType=Unified (>1.5B), # LLM Params=7B2026.03 | 66.6 | |
| UniVideoModel Category=Unified Understanding & Generation Model, Backbone=Qwen-2.5VL-7B2025.10 | 66.6 | |
| Claude3.5-Sonnet2024.09 | 66 | |
| Gemini-1.5-Pro2024.09 | 64 | |
| InternVL2.5Type=Und. Only, LLM Parameters=7B2025.05 | 62.8 | |
| InternVL2.5Type=Und. Only, # LLM Params=7B2026.03 | 62.8 | |
| InternVL3-2BLLM Params=2B, Model Category=Und. Only2026.06 | 62.2 | |
| SPAR-3BLLM Params=2B, Model Category=Und. and Gen.2026.06 | 62.2 | |
| Qwen2-VLType=Und. Only, LLM Parameters=7B2025.05 | 62 | |
| Qwen2-VLType=Und. Only, # LLM Params=7B2026.03 | 62 | |
| Qwen2.5-VLType=Und. Only, LLM Parameters=3B2025.05 | 61.8 | |
| Qwen2.5-VLType=Und. Only, # LLM Params=3B2026.03 | 61.8 | |
| Qwen2.5-VL-3BLLM Params=3B, Model Category=Und. Only2026.06 | 61.8 | |
| OmniGen2Model Category=Unified Understanding & Generation Model2025.10 | 61.8 | |
| InternVL2.5Type=Und. Only, LLM Parameters=1.8B2025.05 | 60.8 | |
| InternVL2.5Type=Und. Only, # LLM Params=1.8B2026.03 | 60.8 | |
| BLIP3-o-4BLLM Params=4B, Model Category=Und. and Gen.2026.06 | 60.1 | |
| DeepSeek-VL2Type=Und. Only, LLM Parameters=4.1B/28B2025.05 | 60 | |
| DeepSeek-VL2Type=Und. Only, # LLM Params=4.1B/28B2026.03 | 60 | |
| SPAR-1BLLM Params=1B, Model Category=Und. and Gen.2026.06 | 59.8 | |
| InternVL3-1BLLM Params=1B, Model Category=Und. Only2026.06 | 59.5 | |
| LLava-OVType=Und. Only, LLM Parameters=7B2025.05 | 57.5 | |
| LLava-OVType=Und. Only, # LLM Params=7B2026.03 | 57.5 | |
| LLaVA-NeXTModel Category=Video Understanding Model2025.10 | 57.4 | |
| Show-o2Model Category=Unified Understanding & Generation Model2025.10 | 56.6 | |
| DPA-Qwen3-32BTarget LLM=Qwen3-32B2026.05 | 56 | |
| MUSE-VLType=Unified, LLM Parameters=32B2025.05 | 55.9 | |
| InternVL2-8B2024.09 | 54.3 | |
| InternVL2Type=Und. Only, LLM Parameters=7B2025.05 | 54.2 | |
| InternVL2Type=Und. Only, # LLM Params=7B2026.03 | 54.2 | |
| Cambrian-34B2024.09 | 53.2 | |
| LVRPOType=Unified, # LLM Params=1.5B2026.03 | 53.2 | |
| POINTS-7BLLM=Qwen-2.5-7B2024.09 | 52.3 | |
| OneVision2024.09 | 51.9 | |
| Ovis1.5-LLAMA3-8B2024.09 | 50.9 | |
| Janus-ProType=Unified, LLM Parameters=7B2025.05 | 50 | |
| POINTS-9BLLM=Yi-1.5-9B2024.09 | 50 | |
| Janus-ProType=Unified (>1.5B), # LLM Params=7B2026.03 | 50 | |
| Janus-Pro-7BLLM Params=7B, Model Category=Und. and Gen.2026.06 | 50 | |
| Qwen2-VLType=Und. Only, LLM Parameters=1.5B2025.05 | 49.5 | |
| Qwen2-VLType=Und. Only, # LLM Params=1.5B2026.03 | 49.5 | |
| Qwen2-VL-2BTarget LLM=Public Models2026.05 | 49.5 | |
| IXC-2.52024.09 | 49.3 | |
| LLaVA-NeXT-Qwen3-32BTarget LLM=Qwen3-32B2026.05 | 49.2 | |
| EMU-2Type=Und. & Gen. Non-prob., LLM=LLaMA-13B, V-Token=CLIP, Resolution=4482024.10 | 48.5 | |
| BAGELType=Unified, LLM Parameters=1.5B MOT2025.05 | 48.2 | |
| BAGELType=Unified, # LLM Params=1.5B2026.03 | 48.2 | |
| BAGEL-7BLLM Params=3B, Model Category=Und. and Gen.2026.06 | 48.2 | |
| TokenFlow-XLModel Category=Unified Understanding & Generation Model2025.10 | 48.2 | |
| InternVL2Type=Und. Only, LLM Parameters=1.8B2025.05 | 44.6 | |
| InternVL2Type=Und. Only, # LLM Params=1.8B2026.03 | 44.6 | |
| Cambrian-1-8BTarget LLM=Public Models2026.05 | 44.2 | |
| SEED-XType=Unified, LLM Parameters=13B2025.05 | 43 | |
| SEED-XType=Unified (>1.5B), # LLM Params=13B2026.03 | 43 | |
| SEED-X-13BLLM Params=13B, Model Category=Und. and Gen.2026.06 | 43 | |
| DPA-Qwen3-4BTarget LLM=Qwen3-4B2026.05 | 42.7 | |
| DPA-LLaMA-3.2-3BTarget LLM=LLaMA-3.2-3B2026.05 | 42 | |
| Idefics3-LLAMA3-8B2024.09 | 41.7 | |
| LLaVA-NeXT-Qwen3-4BTarget LLM=Qwen3-4B2026.05 | 41.4 | |
| LLaVA1.5-13B-BPOModel Size=13B, Baseline Category=reinforcement learning-based2024.11 | 41.4 | |
| TokenFlow-XLType=Unified, LLM Parameters=13B2025.05 | 40.7 | |
| TokenFlow-XLType=Unified (>1.5B), # LLM Params=13B2026.03 | 40.7 | |
| DualTokenLLM=LLaMA-2-7B, Resolution=2562026.03 | 40.5 | |
| EvoTokLLM=Qwen-2.5-7B-Inst., Resolution=2562026.03 | 39.9 | |
| LACINGModel Size=13B2024.11 | 39.9 | |
| Janus-Pro2025.06 | 39.8 | |
| Janus-ProType=Unified, LLM Parameters=1.5B2025.05 | 39.8 | |
| Janus-ProType=Unified, # LLM Params=1.5B2026.03 | 39.8 | |
| LLaVA-NeXT-7BTarget LLM=Public Models2026.05 | 39 | |
| ShareGPT4VLLM=Vicuna-7B, Resolution=3362026.03 | 37.6 | |
| Dynamic-LLaVA-13BILLM=Vicuna-13B, Res.=336, Token=115 (-80%), TFLOPS=4.7+0.01 (-76%)2024.12 | 37.3 | |
| Emu32025.06 | 37.2 | |
| Emu3-ChatType=Und. Only, LLM Parameters=8B2025.05 | 37.2 | |
| EMU-3Type=Joint Prob. Models, LLM=8B from scratch, V-Token=vq-vae, Resolution=5122024.10 | 37.2 | |
| EMU3LLM=8B from scratch, Resolution=-2026.03 | 37.2 | |
| Emu3-ChatType=Und. Only, # LLM Params=8B2026.03 | 37.2 | |
| Emu3-Chat-8BLLM Params=8B, Model Category=Und. Only2026.06 | 37.2 | |
| Emu3Model Category=Unified Understanding & Generation Model2025.10 | 37.2 | |
| Dynamic-LLaVA-TokenPacker-13B_VTProjector=TokenPacker, Image Resolution=336, Prefill Tokens=57, Prefill TFLOPs=2.12024.12 | 37.1 | |
| VIG training + LACINGBase Model=LLaVA-1.5 7B2026.02 | 37.01 | |
| ILLUMEType=Unified, LLM Parameters=7B2025.05 | 37 | |
| ILLUMELLM=Vicuna-7B, Resolution=-2026.03 | 37 | |
| ILLUMEType=Unified (>1.5B), # LLM Params=7B2026.03 | 37 | |
| VCDModel Size=13B, Baseline Category=training-free2024.11 | 36.9 | |
| LLaVA1.5-7B-BPOModel Size=7B, Baseline Category=reinforcement learning-based2024.11 | 36.8 |