Multimodal Understanding on POPE
0.906POPE ScoreInternVL2.5
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| InternVL2.5LLM=InternLM2.5-7B, Token Type=2D-Continuous dynamic2026.05 | 0.906 | — | — | |
| InternVL2.5LLM=InternLM2.5-7B2025.05 | 0.906 | — | — | |
| X-Omni# Tokens=~1T, # Params.=10B / 20B, Architecture Category=Autoregressive meets Diffusion2025.09 | 0.893 | — | — | |
| LLaVA-OVScale=1.5-8B2026.06 | 0.892 | — | — | |
| InternVL3-2BFT-Data=-, Fine-tuned on identical training data=false2026.03 | 0.889 | — | — | |
| LatentUMBasequantized setting=false2026.04 | 0.889 | — | — | |
| ILLUMEType=Und. and Gen., LLM Params=7B2025.01 | 0.885 | — | — | |
| ILLUMELLM=Vicuna-7B, Token Type=2D-Continuous, Res.=2242026.05 | 0.885 | — | — | |
| UnisonStage=Uni. 2 Stage2025.12 | 0.883 | — | — | |
| MOSS-Video-PreviewSFT Mode=real-time2026.06 | 0.8817 | — | — | |
| JanusFlowParams=1.3B, Architecture category=Integrating Autoregressive and Diffusion within Transformers2026.03 | 0.88 | — | — | |
| SpaceR-sft-7BFT-Data=151K, Fine-tuned on identical training data=false2026.03 | 0.88 | — | — | |
| JanusFlow2026.04 | 0.88 | — | — | |
| MOSS-Video-PreviewSFT Mode=offline2026.06 | 0.8788 | — | — | |
| Qwen2.5-VL-7BFT-Data=-, Fine-tuned on identical training data=false2026.03 | 0.878 | — | — | |
| Tokenflow XL2026.04 | 0.878 | — | — | |
| Qwen2.5-VLScale=7B2026.06 | 0.8768 | — | — | |
| Harmo2026.04 | 0.876 | — | — | |
| Qwen2.5-VL-3BFT-Data=-, Fine-tuned on identical training data=false2026.03 | 0.875 | — | — | |
| Janus-Pro-7BType=Und. and Gen., LLM Params=7B2025.01 | 0.874 | — | — | |
| Janus-Pro-7BStage=Uni. 1 Stage2025.12 | 0.874 | — | — | |
| Janus-Pro# Tokens=~300B, # Params.=7B, Architecture Category=Autoregressive w/ Semantic Encoder2025.09 | 0.874 | — | — | |
| Janus-ProParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 0.874 | — | — | |
| Janus Pro2026.04 | 0.874 | — | — | |
| Janus-ProLLM=DeepSeek-LLM-7B, Token Type=2D-Continuous, Res.=3842026.05 | 0.874 | — | — | |
| JanusType=Und. and Gen., # LLM Params=1.3B, External pretrained diffusion model=false2024.10 | 0.87 | — | — | |
| JanusType=Und. and Gen., LLM Params=1.5B2025.01 | 0.87 | — | — | |
| JanusParams=1.5B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 0.87 | — | — | |
| Qwen2.5-VL 7BToken Budget=All 1296 Tokens, Reduction Percentage=0%2026.03 | 0.87 | — | 100 | |
| VG-LLMFT-Data=385k, Fine-tuned on identical training data=false2026.03 | 0.869 | — | — | |
| TokenFlow-XLType=Und. and Gen., LLM Params=13B2025.01 | 0.868 | — | — | |
| TokenFlow-13BStage=Uni. 1 Stage2025.12 | 0.868 | — | — | |
| TokenFlowParams=13B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 0.868 | — | — | |
| Cambrain-S-3BFT-Data=10M, Fine-tuned on identical training data=false2026.03 | 0.868 | — | — | |
| WinTokLLM=Qwen3-8B, Token Type=1D-Continuous, Res.=2562026.05 | 0.865 | — | — | |
| LLaVA-NeXTLLM=Vicuna-7B2025.05 | 0.865 | — | — | |
| WallarooParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction, VLMEvalKit=true2026.03 | 0.864 | — | — | |
| Qwen2.5 VLParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction, VLMEvalKit=true2026.03 | 0.863 | — | — | |
| Janus-Pro-1BType=Und. and Gen., LLM Params=1.5B2025.01 | 0.862 | — | — | |
| Janus-Pro-1BStage=Uni. 1 Stage2025.12 | 0.862 | — | — | |
| LLaVA-v1.5Type=Und. Only, # LLM Params=7B, External pretrained diffusion model=false2024.10 | 0.859 | — | — | |
| LLaVA-v1.5Type=Und. Only, LLM Params=7B2025.01 | 0.859 | — | — | |
| LLaVA-v1.5LLM=Vicuna-7B, Token Type=2D-Continuous, Res.=3362026.05 | 0.859 | — | — | |
| VARGPTLLM=Vicuna-7B, Token Type=2D-Continuous, Res.=2562026.05 | 0.859 | — | — | |
| VILA-UType=Und. and Gen., # LLM Params=7B, External pretrained diffusion model=false2024.10 | 0.858 | — | — | |
| VILA-UType=Und. and Gen., LLM Params=7B2025.01 | 0.858 | — | — | |
| VILA-UStage=Uni. 1 Stage2025.12 | 0.858 | — | — | |
| VILA-UParams=7B, Architecture category=Unifying Understanding and Generation via Next-token Prediction2026.03 | 0.858 | — | — | |
| VILA-U2026.04 | 0.858 | — | — | |
| SpatialLadder-3BFT-Data=26K, Fine-tuned on identical training data=false2026.03 | 0.855 | — | — | |
| LatentUMBasequantized setting=true2026.04 | 0.855 | — | — | |
| GeoSenseFT-Data=940K, Fine-tuned on identical training data=false2026.03 | 0.854 | — | — | |
| Emu3-ChatType=Und. Only, # LLM Params=8B, External pretrained diffusion model=false2024.10 | 0.852 | — | — | |
| Emu3-ChatType=Und. Only, LLM Params=8B2025.01 | 0.852 | — | — | |
| EMU3# Tokens=-, # Params.=8B, Architecture Category=Autoregressive w/o Semantic Encoder2025.09 | 0.852 | — | — | |
| Qwen2.5-VL-3B*FT-Data=940k, Fine-tuned on identical training data=true2026.03 | 0.852 | — | — | |
| EMU32026.04 | 0.852 | — | — | |
| Emu3LLM=8B (from scratch), Token Type=2D-Discrete, Res.=5122026.05 | 0.852 | — | — | |
| VQRAELLM=Vicuna-13B, Token Type=2D-Continuous, Res.=2562026.05 | 0.851 | — | — | |
| LLaVA-PhiType=Und. Only, # LLM Params=2.7B, External pretrained diffusion model=false2024.10 | 0.85 | — | — | |
| LLaVA-PhiType=Und. Only, LLM Params=2.7B2025.01 | 0.85 | — | — | |
| TokenFlow-LLLM=Vicuna-13B, Token Type=2D-Discrete, Res.=2562026.05 | 0.85 | — | — | |
| MobileVLMType=Und. Only, # LLM Params=2.7B, External pretrained diffusion model=false2024.10 | 0.849 | — | — | |
| MobileVLMType=Und. Only, LLM Params=2.7B2025.01 | 0.849 | — | — | |
| TokLIPLLM=Qwen2.5-7B-Instruct, Token Type=1D-Continuous, Res.=3842026.05 | 0.849 | — | — | |
| ViLASR-7BFT-Data=73K, Fine-tuned on identical training data=false2026.03 | 0.848 | — | — | |
| MobileVLM-V2Type=Und. Only, # LLM Params=2.7B, External pretrained diffusion model=false2024.10 | 0.847 | — | — | |
| MobileVLM-V2Type=Und. Only, LLM Params=2.7B2025.01 | 0.847 | — | — | |
| Uni-X# Tokens=240B, # Params.=3B / 4.5B, Architecture Category=Autoregressive w/o Semantic Encoder, Extended Image-Text Training (flag)=true2025.09 | 0.846 | — | — | |
| MobileVLMType=Und. Only, # LLM Params=1.4B, External pretrained diffusion model=false2024.10 | 0.845 | — | — | |
| MobileVLMType=Und. Only, LLM Params=1.4B2025.01 | 0.845 | — | — | |
| Show-o# Tokens=~500B, # Params.=1.3B, Architecture Category=Autoregressive meets Diffusion, Semantic Alignment (flag)=true2025.09 | 0.845 | — | — | |
| VARGPT-9BStage=Uni. 2 Stage2025.12 | 0.844 | — | — | |
| MobileVLM-V2Type=Und. Only, # LLM Params=1.4B, External pretrained diffusion model=false2024.10 | 0.843 | — | — | |
| MobileVLM-V2Type=Und. Only, LLM Params=1.4B2025.01 | 0.843 | — | — | |
| LLaVA-v1.5-Phi-1.5Type=Und. Only, # LLM Params=1.3B, External pretrained diffusion model=false2024.10 | 0.841 | — | — | |
| LLaVA-v1.5-Phi-1.5Type=Und. Only, LLM Params=1.3B2025.01 | 0.841 | — | — | |
| SEED-XLLM=LLaMA2-13B, Token Type=2D-Continuous, Res.=4482026.05 | 0.841 | — | — | |
| WinTokLLM=Qwen3-8B, Token Type=1D-Continuous, Res.=256, semantic tokens=642026.05 | 0.841 | — | — | |
| D-DitType=Und. and Gen., LLM Params=2.0B2025.01 | 0.84 | — | — | |
| VILA-U# Tokens=-, # Params.=7B, Architecture Category=Autoregressive w/ Semantic Encoder2025.09 | 0.839 | — | — | |
| VILA-ULLM=LLaMA2-7B, Token Type=2D-Discrete, Res.=2562026.05 | 0.839 | — | — | |
| Uni-X# Tokens=140B, # Params.=3B / 4.5B, Architecture Category=Autoregressive w/o Semantic Encoder2025.09 | 0.836 | — | — | |
| SemHiTokLLM=Qwen2.5-7B-Instruct, Token Type=2D-Discrete, Res.=2562026.05 | 0.834 | — | — | |
| Liquid# Tokens=-, # Params.=8B, Architecture Category=Autoregressive w/ Semantic Encoder, Semantic Alignment (flag)=true2025.09 | 0.832 | — | — | |
| UniTokLLM=LLaMA2-7B, Token Type=2D-Discrete, Res.=2562026.05 | 0.832 | — | — | |
| Slot-MLLM (7B)LLM=Vicuna-7B2025.05 | 0.83 | — | — | |
| Slot-MLLM (14B)LLM=Qwen2.5-14B-Instruct2025.05 | 0.822 | — | — | |
| PromPruneToken Budget=512, Reduction Percentage=60.5%2026.03 | 0.818 | — | 94.8 | |
| Liquid# Tokens=~90B, # Params.=7B, Architecture Category=Autoregressive w/o Semantic Encoder2025.09 | 0.811 | — | — | |
| LiquidLLM=Gemma-7B, Token Type=2D-Discrete, Res.=5122026.05 | 0.811 | — | — | |
| DivPruneToken Budget=512, Reduction Percentage=60.5%2026.03 | 0.809 | — | 94.2 | |
| PromPruneToken Budget=256, Reduction Percentage=80.2%2026.03 | 0.807 | — | 92.6 | |
| Show-oType=Und. and Gen., LLM Params=1.3B, Resolution=5122025.01 | 0.8 | — | — | |
| Show-oStage=Uni. 1 Stage2025.12 | 0.8 | — | — | |
| Show-o# Tokens=~500B, # Params.=1.3B, Architecture Category=Autoregressive meets Diffusion2025.09 | 0.8 | — | — | |
| Show-oLLM=Phi-1.5-1.3B, Token Type=2D-Discrete, Res.=2562026.05 | 0.8 | — | — | |
| Orthus2026.04 | 0.796 | — | — | |
| InstructBLIPType=Und. Only, # LLM Params=13B, External pretrained diffusion model=false2024.10 | 0.789 | — | — | |
| InstructBLIPType=Und. Only, LLM Params=13B2025.01 | 0.789 | — | — |