Multimodal Model Evaluation on MMBench Chinese
82.6AccuracySAIL-VL2
Evaluation Results
| Method | Links | |
|---|---|---|
| SAIL-VL2Number of Parameters=2.7B2025.12 | 82.6 | |
| CoLLaVONumber of Parameters=7B, Setting=Zero-shot2024.02 | 82.1 | |
| HyperVLNumber of Parameters=1.8B2025.12 | 80.5 | |
| SAIL-VL1.5Number of Parameters=2.5B2025.12 | 79.3 | |
| HyperVL ViTLNumber of Parameters=2.0B2025.12 | 79.3 | |
| Ovis2Number of Parameters=2.5B2025.12 | 78.9 | |
| AndesVLNumber of Parameters=2.4B2025.12 | 78.6 | |
| InternVL3Number of Parameters=2.1B2025.12 | 78.3 | |
| Qwen3-VLNumber of Parameters=2.1B2025.12 | 78 | |
| InternVL3.5Number of Parameters=2.3B2025.12 | 77.7 | |
| MoAIZero-shot=true, Number of Parameters=7B2024.03 | 76.5 | |
| Qwen2-VLNumber of Parameters=2.2B2025.12 | 74.2 | |
| LLaVA-XTunerNumber of Parameters=20B, Setting=Zero-shot2024.02 | 73.7 | |
| LLaVA-XTunerZero-shot=true, Number of Parameters=20B2024.03 | 73.7 | |
| Intern-XCNumber of Parameters=7B, Setting=Zero-shot2024.02 | 72.4 | |
| Intern-XCZero-shot=true, Number of Parameters=7B2024.03 | 72.4 | |
| VisionLLM v2-chatBackbone=Vicuna-7B, #Data=22M2024.06 | 67.6 | |
| Dragonfly (Ours)Backbone=Llama3-8B, #Data=2.9M2024.06 | 66.1 | |
| Merlin2023.11 | 65.5 | |
| VILA-13B + ShareGPT4VLLM=Llama-2-13B, Resolution=336, Pre-training samples=50M, Instruction tuning samples=1M2023.12 | 65.4 | |
| FMVR-LLaVA#Vision Tokens=576, Backbone=LLaVA-1.5-13B2026.03 | 65.1 | |
| LLaVA-UHD#Data=1.2M, MaxRes.=672x1008, AR.=Any, TFLOPS=14.62024.03 | 64.8 | |
| VILA-13BLLM=Llama-2-13B, Resolution=336, Pre-training samples=50M, Instruction tuning samples=1M2023.12 | 64.3 | |
| ASMv2Model Scale=13B2024.02 | 64.3 | |
| FMVR-LLaVA#Vision Tokens=144, Backbone=LLaVA-1.5-13B2026.03 | 63.8 | |
| LLaVA-1.5#Data=1.2M, MaxRes.=336x336, AR.=Fix, TFLOPS=15.52024.03 | 63.6 | |
| LLaVA-1.5LLM=Vicuna-1.5-13B, Resolution=336, Pre-training samples=0.6M, Instruction tuning samples=0.7M2023.12 | 63.6 | |
| LLaVA1.5Number of Parameters=13B, Setting=Zero-shot2024.02 | 63.6 | |
| LLaVA-v1.5-13BLLM=Vicuna-13B, Resolution=3362024.05 | 63.6 | |
| LLaVA-1.5LLM=Vicuna-13B, Data=558K+665K2024.01 | 63.6 | |
| CaMML-13BLLM=Vicuna-13B, Data=558K++665K2024.01 | 63.6 | |
| LLaVA1.5Zero-shot=true, Number of Parameters=13B2024.03 | 63.6 | |
| LLaVA-1.5Model Scale=13B2024.02 | 63.6 | |
| LLaVA-1.5-13B#Vision Tokens=576, Backbone=LLaVA-1.5-13B2026.03 | 63.5 | |
| LLaVA-1.5-13BResolution=336*3362024.11 | 63.3 | |
| VILAModel Scale=13B2024.02 | 62.8 | |
| PyramidDrop#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 62.8 | |
| SparseVLM#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 62.6 | |
| VisionZip#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 62.5 | |
| FMVR-LLaVA#Vision Tokens=36, Backbone=LLaVA-1.5-13B2026.03 | 62.5 | |
| FastV#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 62.3 | |
| ShareGPT4VNumber of Parameters=7B, Setting=Zero-shot2024.02 | 62.2 | |
| ShareGPT4VZero-shot=true, Number of Parameters=7B2024.03 | 62.2 | |
| ShareGPT4V-7BBackbone=Vicuna-7B, #Data=1.8M2024.06 | 62.2 | |
| DART#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 62.2 | |
| PPLLAVAResolution=336*3362024.11 | 62 | |
| Qwen-VL-Chat2023.11 | 61.8 | |
| VILA-7BLLM=Llama-2-7B, Resolution=336, Pre-training samples=50M, Instruction tuning samples=1M2023.12 | 61.7 | |
| VILABackbone=Llama2-7B, #Data=61M2024.06 | 61.7 | |
| CDPruner#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 61.5 | |
| VisPruner#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 61.3 | |
| Prumerge+#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 61.2 | |
| MQT-LLaVA#Vision Tokens=144, Backbone=LLaVA-1.5-13B2026.03 | 61 | |
| DivPrune#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 60.7 | |
| HyperLLaVALLM=Vicuna-7B, Res.=336, PT=558K, IT=665K2024.03 | 60.6 | |
| CaMML-7BLLM=Vicuna-7B, Data=558K++665K2024.01 | 60.6 | |
| LLaVA-NeXTBackbone=Vicuna-7B, #Data=1.2M2024.06 | 60.6 | |
| LLaVA-Next-7BResolution=672*6722024.11 | 60.6 | |
| FMVR-LLaVA#Vision Tokens=9, Backbone=LLaVA-1.5-13B2026.03 | 60.4 | |
| Shikra2023.11 | 60.2 | |
| M3#Vision Tokens=144, Backbone=LLaVA-1.5-13B2026.03 | 59.8 | |
| HyperLLaVA w/o EvLLM=Vicuna-7B, Res.=336, PT=558K, IT=665K2024.03 | 59.7 | |
| LLaVA-1.52023.11 | 59.5 | |
| InfMLLM-7B-ChatLLM=Vicuna-7B2023.11 | 59.1 | |
| HyperLLaVA w/o ELLLM=Vicuna-7B, Res.=336, PT=558K, IT=665K2024.03 | 58.6 | |
| TRIM#Vision Tokens=128, Backbone=LLaVA-1.5-13B2026.03 | 58.4 | |
| LLaVA-1.5-7BLLM=Vicuna-7B2023.11 | 58.3 | |
| LLaVA-1.5LLM=Vicuna-7B, Res.=336, PT=558K, IT=665K2024.03 | 58.3 | |
| LLaVA-1.5LLM=Vicuna-1.5-7B, Resolution=336, Pre-training samples=0.6M, Instruction tuning samples=0.7M2023.12 | 58.3 | |
| LLaVA1.5Number of Parameters=7B, Setting=Zero-shot2024.02 | 58.3 | |
| LLaVA-1.5LLM=Vicuna-7B, Data=558K+665K2024.01 | 58.3 | |
| LLaVA1.5Zero-shot=true, Number of Parameters=7B2024.03 | 58.3 | |
| LLaVA-1.5Backbone=Vicuna-7B, #Data=1.2M2024.06 | 58.3 | |
| FSRRetained Tokens=960, Reduction Ratio=↓ 66.7%2026.02 | 58.3 | |
| VisPrunerRetained Tokens=960, Reduction Ratio=↓ 66.7%2026.02 | 58.2 | |
| SPHINX-2k#Data=1.0B, MaxRes.=762x762, AR.=Fix, TFLOPS=69.42024.03 | 57.9 | |
| SPHINX-2kBackbone=Llama2-7B, #Data=1B2024.06 | 57.9 | |
| FSRRetained Tokens=640, Reduction Ratio=↓ 77.8%2026.02 | 57.9 | |
| InternVL-7BBackbone=Vicuna-7B, #Data=>28.7B2024.06 | 57.6 | |
| CDPrunerRetained Tokens=960, Reduction Ratio=↓ 66.7%2026.02 | 57.6 | |
| CDPrunerRetained Tokens=640, Reduction Ratio=↓ 77.8%2026.02 | 57.6 | |
| LLaVA-NeXT-7BRetained Tokens=2880, Reduction Ratio=100%2026.02 | 57.4 | |
| DivPruneRetained Tokens=640, Reduction Ratio=↓ 77.8%2026.02 | 57.3 | |
| VisPrunerRetained Tokens=640, Reduction Ratio=↓ 77.8%2026.02 | 57.3 | |
| LLaVA-NeXT-7B#Vision Tokens=28802026.03 | 57.3 | |
| FMVR-LLaVA#Vision Tokens=7202026.03 | 57.1 | |
| HoloVRetained Tokens=320, Reduction Ratio=↓ 88.9%2026.02 | 57 | |
| FMVR-LLaVA#Vision Tokens=28802026.03 | 56.9 | |
| Qwen-VL-ChatLLM=Qwen-7B2023.11 | 56.7 | |
| Qwen-VL-ChatLLM=Qwen-7B, Res.=448, PT=1.4B, IT=50M2024.03 | 56.7 | |
| Qwen-VL-ChatLLM=Qwen-7B, Resolution=448, Pre-training samples=1.4B, Instruction tuning samples=50M2023.12 | 56.7 | |
| Qwen-VL-ChatNumber of Parameters=7B, Setting=Zero-shot2024.02 | 56.7 | |
| Qwen-VL-ChatLLM=Qwen-7B, Data=1.4B+50M2024.01 | 56.7 | |
| Qwen-VL-ChatZero-shot=true, Number of Parameters=7B2024.03 | 56.7 | |
| Qwen-VL-Chat2024.02 | 56.7 | |
| Qwen-VL-ChatBackbone=Qwen-7B, #Data=1.4B2024.06 | 56.7 | |
| LLaVA-Next-VideoResolution=336*3362024.11 | 56.7 | |
| HoloVRetained Tokens=640, Reduction Ratio=↓ 77.8%2026.02 | 56.7 | |
| SparseVLM#Vision Tokens=3202026.03 | 56.7 | |
| FMVR-LLaVA#Vision Tokens=1802026.03 | 56.3 |