Multimodal Perception on MME Perception
1,748.4Perception ScoreInternVL3-8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVL3-8BModel Category=MMU-Experts2026.03 | 1,748.4 | — | |
| Dynin-OmniModel Category=Unified, Native speech capability=true2026.03 | 1,733.6 | — | |
| Qwen2.5-VL-7BModel Category=MMU-Experts2026.03 | 1,698.1 | — | |
| CoLLaVONumber of Parameters=7B, Setting=Zero-shot2024.02 | 1,689.7 | — | |
| BAGELModel Category=Unified2026.03 | 1,687 | — | |
| OmniVinciModel Category=Perception-centric2026.03 | 1,651 | — | |
| MUSE-3B# LLM Params=2B2026.05 | 1,645 | — | |
| UniLIP-3B# LLM Params=2B2026.05 | 1,636 | — | |
| InternVL3# LLM Params=1.8B2026.05 | 1,633 | — | |
| Baichuan-Omni-1.5Model Category=Perception-centric2026.03 | 1,632.1 | — | |
| Ovis2-8BModel Category=MMU-Experts2026.03 | 1,628.1 | — | |
| Show-o2Model Category=Unified, Video generation support=true2026.03 | 1,620.5 | — | |
| BAGEL# LLM Params=3B2026.05 | 1,610 | — | |
| Tar# LLM Params=7B2026.05 | 1,571 | — | |
| ShareGPT4VNumber of Parameters=7B, Setting=Zero-shot2024.02 | 1,567.4 | — | |
| ShareGPT4VLLM=Vicuna-7B, Category=Understanding Only2024.12 | 1,567.4 | — | |
| Janus-ProModel Category=Unified2026.03 | 1,567.1 | — | |
| Janus-Pro# LLM Params=7B2026.05 | 1,567 | — | |
| LLaVA-1.5-13BLLM=Vicuna-1.5-13B, Visual Encoder=CLIP-L@336, #PT=558K, #FT=665K2024.05 | 1,541.7 | — | |
| NExT-OMNIModel Category=Unified, Native speech capability=true, Video generation support=true2026.03 | 1,537.8 | — | |
| OpenOmniModel Category=Perception-centric2026.03 | 1,536.9 | — | |
| Lumina-DiMOOModel Category=Unified2026.03 | 1,534.2 | — | |
| LLaVA1.5Number of Parameters=13B, Setting=Zero-shot2024.02 | 1,531.3 | — | |
| Intern-XCNumber of Parameters=7B, Setting=Zero-shot2024.02 | 1,528.4 | — | |
| BLIP3-o# LLM Params=4B2026.05 | 1,528 | — | |
| LLaVA1.5Number of Parameters=7B, Setting=Zero-shot2024.02 | 1,510.7 | — | |
| LLaVA-v1.52024.10 | 1,510.7 | — | |
| LLaVA-1.5LLM=Vicuna-7B, Category=Understanding Only2024.12 | 1,510.7 | — | |
| Imp-4BLLM=Phi-3-3.8B, Visual Encoder=SigLIP-SO@384, #PT=558K, #FT=1M2024.05 | 1,507.7 | — | |
| LLaDA-VModel Category=MMU-Experts2026.03 | 1,507 | — | |
| LLaVA-1.5-7BSparsity Op.=false, Sparsity Tok.=false, Vision-side Token=576, Vision-side TFLOPs (Rel.)=7.65 (100.0%)2026.02 | 1,506.5 | — | |
| MUSE-1B# LLM Params=1B2026.05 | 1,505 | — | |
| HA-DPOBase Model=LLaVA-v1.52024.10 | 1,502.6 | — | |
| UniLIP-1B# LLM Params=1B2026.05 | 1,499 | — | |
| Bunny-4BLLM=Phi-3-3.8B, Visual Encoder=SigLIP-SO@384, #PT=2M, #FT=695K2024.05 | 1,495.2 | — | |
| InternVL3# LLM Params=1B2026.05 | 1,492 | — | |
| PDropSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=270, Vision-side TFLOPs (Rel.)=3.54 (46.3%)2026.02 | 1,490.1 | — | |
| Bunny-3BLLM=Phi-2-2.7B, Visual Encoder=SigLIP-SO@384, #PT=2M, #FT=695K2024.05 | 1,488.8 | — | |
| Qwen-VL-ChatNumber of Parameters=7B, Setting=Zero-shot2024.02 | 1,487.5 | — | |
| Qwen-VL-ChatLLM=Qwen-7B, Category=Understanding Only2024.12 | 1,487.5 | — | |
| FudokiModel Category=Unified2026.03 | 1,485.4 | — | |
| Qwen2.5-Omni-7BModel Category=Perception-centric2026.03 | 1,481.3 | — | |
| Dynamic-LLaVASparsity Op.=false, Sparsity Tok.=true, Vision-side Token=115, Vision-side TFLOPs (Rel.)=1.50 (19.6%)2026.02 | 1,479.8 | — | |
| LLaVA-1.5-7BLLM=Vicuna-1.5-7B, Visual Encoder=CLIP-L@336, #PT=558K, #FT=665K2024.05 | 1,476.9 | — | |
| ViCASparsity Op.=false, Sparsity Tok.=true, Vision-side Token=24, Vision-side TFLOPs (Rel.)=0.31 (4.1%)2026.02 | 1,464.5 | — | |
| SEED-X# LLM Params=13B2026.05 | 1,457 | — | |
| MiniCPM-V-3BLLM=MiniCPM-SFT-2B, Visual Encoder=SigLIP-SO@384, #PT=300M, #FT=8M2024.05 | 1,452 | — | |
| SilkieRTBase Model=Qwen-VL-Chat, DPO Training Subset=VLFeedback Red Teaming Subset2024.10 | 1,450.9 | — | |
| mPLUG-Owl2Number of Parameters=7B, Setting=Zero-shot2024.02 | 1,450.2 | — | |
| mPLUG-Owl2-8BLLM=LLAMA2-7B, Visual Encoder=CLIP-L@448, #PT=400M, #FT=1.2M2024.05 | 1,450.2 | — | |
| ViCA+PDrop†Sparsity Op.=true, Sparsity Tok.=true, Vision-side Token=12, Vision-side TFLOPs (Rel.)=0.16 (2.0%)2026.02 | 1,449.7 | — | |
| Imp-3BLLM=Phi-2-2.7B, Visual Encoder=SigLIP-SO@384, #PT=558K, #FT=1M2024.05 | 1,446.4 | — | |
| ILLUMELLM=Vicuna-7B, Category=Unify Understanding and Generation2024.12 | 1,445.3 | — | |
| ILLUME# LLM Params=7B2026.05 | 1,445 | — | |
| Qwen-VL-Chat2024.10 | 1,439.1 | — | |
| SEED-XModel Category=Unified2026.03 | 1,435.7 | — | |
| POVIDBase Model=LLaVA-v1.52024.10 | 1,423.9 | — | |
| YOPOSparsity Op.=true, Sparsity Tok.=false, Vision-side Token=70, Vision-side TFLOPs (Rel.)=0.92 (12.0%)2026.02 | 1,423.5 | — | |
| VISASparsity Op.=false, Sparsity Tok.=true, Vision-side Token=64, Vision-side TFLOPs (Rel.)=0.83 (10.9%)2026.02 | 1,420.6 | — | |
| TRIMSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=29, Vision-side TFLOPs (Rel.)=0.38 (4.9%)2026.02 | 1,415.4 | — | |
| MMaDAModel Category=Unified2026.03 | 1,410.7 | — | |
| TokLIP# LLM Params=7B2026.05 | 1,410 | — | |
| TwigVLMSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=64, Vision-side TFLOPs (Rel.)=0.83 (10.9%)2026.02 | 1,404 | — | |
| VILA-ULLM=LLaMA-2-7B, Input Resolution=384, Category=Unify Understanding and Generation2024.12 | 1,401.8 | — | |
| DOPCDSparsity Op.=true, Sparsity Tok.=true, Vision-side Token=32, Vision-side TFLOPs (Rel.)=0.42 (5.4%)2026.02 | 1,397.5 | — | |
| TokenPackerSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=16, Vision-side TFLOPs (Rel.)=0.21 (2.7%)2026.02 | 1,378.8 | — | |
| Delta-LLaVASparsity Op.=false, Sparsity Tok.=true, Vision-side Token=16, Vision-side TFLOPs (Rel.)=0.21 (2.7%)2026.02 | 1,375.9 | — | |
| CDPrunerSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=32, Vision-side TFLOPs (Rel.)=0.42 (5.4%)2026.02 | 1,373 | — | |
| TokenFlow-B# LLM Params=13B2026.05 | 1,354 | — | |
| LLaVA-PruMergeSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=32, Vision-side TFLOPs (Rel.)=0.42 (5.4%)2026.02 | 1,350.3 | — | |
| Mini-Gemini-2BLLM=Gemma-2B, Visual Encoder=CLIP-L@336, #PT=1.2M, #FT=1.5M2024.05 | 1,341 | — | |
| JanusLLM=DeepSeek-LLM-1.3B, Category=Unify Understanding and Generation2024.12 | 1,338 | — | |
| VILA-ULLM=LLaMA-2-7B, Input Resolution=256, Category=Unify Understanding and Generation2024.12 | 1,336.2 | — | |
| VILA-U# LLM Params=7B2026.05 | 1,336 | — | |
| LLaVA-Phi-3BLLM=Phi-2-2.7B, Visual Encoder=CLIP-L@336, #PT=558K, #FT=665K2024.05 | 1,335.1 | — | |
| LLaVA-SFT2024.10 | 1,315.7 | — | |
| PDropSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=64, Vision-side TFLOPs (Rel.)=0.83 (10.9%)2026.02 | 1,309.2 | — | |
| HyperCLOVAX-8B-OmniModel Category=Unified, Native speech capability=true2026.03 | 1,307 | — | |
| DOPvSparsity Op.=true, Sparsity Tok.=true, Vision-side Token=32, Vision-side TFLOPs (Rel.)=0.42 (5.4%)2026.02 | 1,306.5 | — | |
| Imp-2BLLM=Qwen-1.5-1.8B, Visual Encoder=SigLIP-SO@384, #PT=558K, #FT=1M2024.05 | 1,304.8 | — | |
| Bunny-2BLLM=Qwen1.5-1.8B, Visual Encoder=SigLIP-SO@384, #PT=2M, #FT=695K2024.05 | 1,300.8 | — | |
| BLIP2Number of Parameters=13B, Setting=Zero-shot2024.02 | 1,293.8 | — | |
| OtterNumber of Parameters=7B, Setting=Zero-shot2024.02 | 1,292.3 | — | |
| MobileVLM-3BLLM=MobileLLaMA-2.7B, Visual Encoder=CLIP-L@336, #PT=558K, #FT=665K2024.05 | 1,288.9 | — | |
| DivPruneSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=32, Vision-side TFLOPs (Rel.)=0.42 (5.4%)2026.02 | 1,284.9 | — | |
| VisPrunerSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=32, Vision-side TFLOPs (Rel.)=0.42 (5.4%)2026.02 | 1,271 | — | |
| SPHINX-X-8BLLM=InternLM2-7B, Visual Encoder=CLIP&Dino@224, #PT=0, #FT=15M2024.05 | 1,260.4 | — | |
| Emu3-Chat# LLM Params=8B2026.05 | 1,244 | — | |
| LLaVA-OV# LLM Params=1B2026.05 | 1,238 | — | |
| MetaQuery-B# LLM Params=1B2026.05 | 1,238 | — | |
| InstructBLIPNumber of Parameters=13B, Setting=Zero-shot2024.02 | 1,212.8 | — | |
| LLaVA-RLHFBase Model=LLaVA-SFT2024.10 | 1,203.3 | — | |
| Harmon# LLM Params=1.5B2026.05 | 1,155 | — | |
| Show-O# LLM Params=1.3B2026.05 | 1,097 | — | |
| SparseVLMSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=32, Vision-side TFLOPs (Rel.)=0.42 (5.4%)2026.02 | 1,046.7 | — | |
| mPLUG-OwlNumber of Parameters=7B, Setting=Zero-shot2024.02 | 967.3 | — | |
| Show-oLLM=Phi-1.5B, Category=Unify Understanding and Generation2024.12 | 948.4 | — | |
| FastVSparsity Op.=false, Sparsity Tok.=true, Vision-side Token=32, Vision-side TFLOPs (Rel.)=0.42 (5.4%)2026.02 | 884.6 | — | |
| LLaVANumber of Parameters=7B, Setting=Zero-shot2024.02 | 807 | — | |
| BLIP-2LLM Backbone=FLAN-T52023.12 | — | 1,293.8 |