Multimodal Benchmarking on MMBench (Accuracy)
87.27AccuracyVISE
Evaluation Results
| Method | Links | |
|---|---|---|
| VISEModel Scale=32B2026.06 | 87.27 | |
| VisionZero-CLEVRModel Scale=32B2026.06 | 87.18 | |
| VisionZero-ChartModel Scale=32B2026.06 | 87.13 | |
| iReasonerModel Scale=32B2026.06 | 87.08 | |
| EvoLMMModel Scale=32B2026.06 | 87.02 | |
| VisionZero-RWModel Scale=32B2026.06 | 86.72 | |
| BaseModel Scale=32B2026.06 | 86.65 | |
| VisPlayModel Scale=32B2026.06 | 86.58 | |
| VISEModel Scale=8B2026.06 | 85.46 | |
| VisionZero-CLEVRModel Scale=8B2026.06 | 85.11 | |
| iReasonerModel Scale=8B2026.06 | 85.09 | |
| VisionZero-ChartModel Scale=8B2026.06 | 85.07 | |
| EvoLMMModel Scale=8B2026.06 | 85.02 | |
| VisionZero-CLEVRModel Scale=4B2026.06 | 84.83 | |
| VisionZero-RWModel Scale=8B2026.06 | 84.76 | |
| BaseModel Scale=8B2026.06 | 84.71 | |
| VisPlayModel Scale=8B2026.06 | 84.62 | |
| VISEModel Scale=4B2026.06 | 84.51 | |
| InternVL2.5-8BBase Model=InternVL2.5-8B, Token Reduction Ratio=0%2026.04 | 84.4 | |
| iReasonerModel Scale=4B2026.06 | 84.34 | |
| EvoLMMModel Scale=4B2026.06 | 84.21 | |
| Qwen2.5-VL-7BBase Model=Qwen2.5-VL-7B, Token Reduction Ratio=0%2026.04 | 83.9 | |
| VisionZero-ChartModel Scale=4B2026.06 | 83.78 | |
| VanillaBackbone=Qwen3-VL-4B-FP8, Flops Ratio Reduction=0%, Precision=FP82026.04 | 83.6 | |
| VisPlayModel Scale=4B2026.06 | 83.47 | |
| BaseModel Scale=4B2026.06 | 83.42 | |
| VisionZero-RWModel Scale=4B2026.06 | 83.32 | |
| VanillaBackbone=Qwen2-VL-7B, Flops Ratio Reduction=0%2026.04 | 82.5 | |
| ID-SelectionBase Model=Qwen2.5-VL-7B, Token Reduction Ratio=77.8%, Importance Source=LLM second-layer cross-modal attention2026.04 | 82.1 | |
| ID-SelectionBase Model=InternVL2.5-8B, Token Reduction Ratio=77.8%, Importance Source=LLM second-layer cross-modal attention2026.04 | 80.9 | |
| HalfVBackbone=Qwen3-VL-4B-FP8, Flops Ratio Reduction=77.8%, Precision=FP82026.04 | 80.9 | |
| FastVBase Model=Qwen2.5-VL-7B, Token Reduction Ratio=77.8%2026.04 | 80.6 | |
| FastVBase Model=InternVL2.5-8B, Token Reduction Ratio=77.8%2026.04 | 80.5 | |
| ID-SelectionBase Model=Qwen2.5-VL-7B, Token Reduction Ratio=88.9%, Importance Source=LLM second-layer cross-modal attention2026.04 | 80 | |
| DARTBase Model=InternVL2.5-8B, Token Reduction Ratio=77.8%2026.04 | 80 | |
| DARTBase Model=Qwen2.5-VL-7B, Token Reduction Ratio=77.8%2026.04 | 79.8 | |
| MM1LLM Backbone=MM1-7B-MoE2026.05 | 79.7 | |
| HalfVBackbone=Qwen2-VL-7B, Flops Ratio Reduction=88.9%2026.04 | 79.2 | |
| DARTBackbone=Qwen3-VL-4B-FP8, Flops Ratio Reduction=77.8%, Precision=FP82026.04 | 79.2 | |
| ID-SelectionBase Model=InternVL2.5-8B, Token Reduction Ratio=88.9%, Importance Source=LLM second-layer cross-modal attention2026.04 | 78.8 | |
| FastVBackbone=Qwen3-VL-4B-FP8, Flops Ratio Reduction=77.8%, Precision=FP82026.04 | 77.5 | |
| HalfVBackbone=Qwen3-VL-4B-FP8, Flops Ratio Reduction=88.9%, Precision=FP82026.04 | 77.4 | |
| AsyMoE-LLaMA3LLM Backbone=LLaMA-3-8B, Act Param=10.8B2026.05 | 76.8 | |
| VISEModel Scale=2B2026.06 | 76.72 | |
| DARTBackbone=Qwen3-VL-4B-FP8, Flops Ratio Reduction=88.9%, Precision=FP82026.04 | 76.2 | |
| AsyMoE-Phi3LLM Backbone=Phi-3-mini, Act Param=4.1B2026.05 | 75.9 | |
| FastVBase Model=InternVL2.5-8B, Token Reduction Ratio=88.9%2026.04 | 75.8 | |
| MoIIE-LLaMA3LLM Backbone=LLaMA-3-8B, Act Param=11.3B2026.05 | 75.7 | |
| Mini-GeminiLLM Backbone=Mixtral-8×7B, Act Param=13.5B2026.05 | 75.6 | |
| MoIIE-Phi3LLM Backbone=Phi-3-mini, Act Param=5.5B2026.05 | 75.4 | |
| CuMoLLM Backbone=Mixtral-8×7B, Act Param=13.5B2026.05 | 75.3 | |
| VisionZero-CLEVRModel Scale=2B2026.06 | 75.07 | |
| DARTBase Model=InternVL2.5-8B, Token Reduction Ratio=88.9%2026.04 | 75 | |
| FastVBackbone=Qwen3-VL-4B-FP8, Flops Ratio Reduction=88.9%, Precision=FP82026.04 | 74.9 | |
| ContextMoE-Phi2LLM Backbone=Phi-2, Act Param=4.1B2026.05 | 74.9 | |
| DARTBackbone=Qwen2-VL-7B, Flops Ratio Reduction=88.9%2026.04 | 74.8 | |
| iReasonerModel Scale=2B2026.06 | 74.75 | |
| EvoLMMModel Scale=2B2026.06 | 74.62 | |
| VisPlayModel Scale=2B2026.06 | 74.52 | |
| VisionZero-ChartModel Scale=2B2026.06 | 74.51 | |
| BaseModel Scale=2B2026.06 | 74.48 | |
| VisionZero-RWModel Scale=2B2026.06 | 74.4 | |
| DARTBase Model=Qwen2.5-VL-7B, Token Reduction Ratio=88.9%2026.04 | 74.1 | |
| FastVBase Model=Qwen2.5-VL-7B, Token Reduction Ratio=88.9%2026.04 | 73.8 | |
| CapRL-Image-5MModel Configuration=InternLM2.5-7B + CLIP-ViT-L2026.06 | 73.7 | |
| CapRL-Image-1MModel Configuration=InternLM2.5-7B + CLIP-ViT-L2026.06 | 73.4 | |
| Deepseek-VLLLM Backbone=Deepseek-7B2026.05 | 73.2 | |
| CapRL-Image-5MModel Configuration=Qwen2.5-3B + Qwen2.5-ViT2026.06 | 73.1 | |
| CapRL-Image-5MModel Configuration=Qwen2.5-7B + Qwen2.5-ViT2026.06 | 72.7 | |
| DenseFusion-1MModel Configuration=Qwen2.5-7B + Qwen2.5-ViT2026.06 | 72.6 | |
| ShareGPT4V-1MModel Configuration=InternLM2.5-7B + CLIP-ViT-L2026.06 | 72.6 | |
| HoloVBackbone=Qwen2-VL-7B, Flops Ratio Reduction=88.9%2026.04 | 72.4 | |
| LLaVA-LLaMA3LLM Backbone=LLaMA-3-8B, Act Param=8.4B2026.05 | 72.3 | |
| DenseFusion-1MModel Configuration=InternLM2.5-7B + CLIP-ViT-L2026.06 | 72.2 | |
| VanillaModel Configuration=Qwen2.5-7B + Qwen2.5-ViT2026.06 | 72.1 | |
| CapRL-Image-1MModel Configuration=Qwen2.5-7B + Qwen2.5-ViT2026.06 | 72.1 | |
| ShareGPT4V-1MModel Configuration=Qwen2.5-7B + Qwen2.5-ViT2026.06 | 71.9 | |
| SPHINX-XLLM Backbone=Mixtral-8×7B2026.05 | 71.3 | |
| VanillaModel Configuration=InternLM2.5-7B + CLIP-ViT-L2026.06 | 70.7 | |
| CapRL-Image-1MModel Configuration=Qwen2.5-3B + Qwen2.5-ViT2026.06 | 70.5 | |
| LLaVA-NeXT Vanilla 13BToken Budget=2880, Avg. Images=4.12, Avg. Tokens=2373.122025.08 | 70 | |
| Mini-GeminiLLM Backbone=Vicuna-7B, Act Param=7.3B2026.05 | 69.3 | |
| FastVBackbone=Qwen2-VL-7B, Flops Ratio Reduction=88.9%2026.04 | 69.2 | |
| DenseFusion-1MModel Configuration=Qwen2.5-3B + Qwen2.5-ViT2026.06 | 69 | |
| ShareGPT4V-1MModel Configuration=Qwen2.5-3B + Qwen2.5-ViT2026.06 | 68.9 | |
| ShareGPT-4VLLM Backbone=Vicuna-7B2026.05 | 68.8 | |
| MousiLLM Backbone=Vicuna-7B, Act Param=7.9B2026.05 | 68.8 | |
| VisionZipToken Budget=640, Avg. Tokens=5272025.08 | 68.6 | |
| VanillaModel Configuration=Qwen2.5-3B + Qwen2.5-ViT2026.06 | 68.6 | |
| MoE-LLaVA-Phi2LLM Backbone=Phi-2, Act Param=3.6B2026.05 | 68 | |
| MMTokToken Budget=640, Avg. Tokens=5272025.08 | 67.44 | |
| LLaVA-NeXT-7BBase Model=LLaVA-NeXT-7B, Retained Tokens=2880, Pruning Ratio=100%2026.04 | 67.4 | |
| LLaVA-NeXTLLM Backbone=Vicuna-7B, Act Param=7.1B2026.05 | 67.4 | |
| VisionZipToken Budget=320, Avg. Tokens=2642025.08 | 67.2 | |
| HiPruneRetained Tokens=640, Token Reduction Ratio=↓ 77.8%2026.07 | 67 | |
| VisionZip ✨Token Budget=320, Avg. Tokens=2642025.08 | 66.9 | |
| DivPruneToken Budget=640, Avg. Tokens=5272025.08 | 66.84 | |
| VisionZip ✨Token Budget=640, Avg. Tokens=5272025.08 | 66.6 | |
| EADPRetained Tokens=640, Token Reduction Ratio=↓ 77.8%2026.07 | 66.5 | |
| CDPrunerRetained Tokens=640, Token Reduction Ratio=↓ 77.8%2026.07 | 66.2 |