Multi-modal Question Answering on MMBench
86.4AccuracyQwen2.5vl-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5vl-InstructModel Scale=7B, Training Protocol=PSO, Category=Open-Source Reasoning MLLMs2025.12 | 86.4 | |
| Qwen2.5vl-InstructModel Scale=7B, Training Protocol=SFT + GRPO, Category=Open-Source Reasoning MLLMs2025.12 | 85.1 | |
| MergeMixbaseline=SFT Vision2025.10 | 84.19 | |
| Qwen2.5-VL-Ins-7B2025.10 | 84.02 | |
| Qwen3-VL-4BModel Scale=4B2026.02 | 83.9 | |
| DenseMLLM-4BModel Scale=4B2026.02 | 83.9 | |
| InternVL2.5-8B-VisualPRMModel Scale=8B, Category=Open-Source Reasoning MLLMs2025.12 | 83.5 | |
| SFT Vision2025.10 | 83.41 | |
| Qwen2.5vl-InstructModel Scale=7B, Category=Open-Source Reasoning MLLMs2025.12 | 83.3 | |
| Qwen2.5-VL-7BToken Budget=100%2025.08 | 83.2 | |
| VisionThink-7B2025.10 | 82.73 | |
| HiPruneBase Model=Qwen2.5-VL-7B, Token Budget=33.3%2025.08 | 82.6 | |
| HiPrune++Base Model=Qwen2.5-VL-7B, Token Budget=33.3%2025.08 | 82.1 | |
| Cambrian-1Model Scale=34B, Category=Open-Source General MLLMs2025.12 | 81.4 | |
| InternVL-3.5-4BModel Scale=4B2026.02 | 80.3 | |
| HiPrune++Base Model=Qwen2.5-VL-7B, Token Budget=22.2%2025.08 | 80.3 | |
| HiPruneBase Model=Qwen2.5-VL-7B, Token Budget=22.2%2025.08 | 80.2 | |
| Curr-ReFTModel Scale=7B, Category=Open-Source Reasoning MLLMs2025.12 | 79 | |
| Qwen2.5-VL-3BToken Budget=100%2025.08 | 77.3 | |
| HiPruneBase Model=Qwen2.5-VL-7B, Token Budget=11.1%2025.08 | 76.1 | |
| HiPruneBase Model=Qwen2.5-VL-3B, Token Budget=33.3%2025.08 | 75.9 | |
| R1-OnevisionModel Scale=7B, Category=Open-Source Reasoning MLLMs2025.12 | 75.6 | |
| HiPrune++Base Model=Qwen2.5-VL-3B, Token Budget=33.3%2025.08 | 75.6 | |
| HiPrune++Base Model=Qwen2.5-VL-7B, Token Budget=11.1%2025.08 | 75.6 | |
| AvgModel size=7B, Evaluation protocol=0-shot2026.02 | 75.52 | |
| GPT-4VCategory=Open-Source General MLLMs2025.12 | 75 | |
| MaD-MixModel size=7B, Evaluation protocol=0-shot2026.02 | 74.57 | |
| HiPruneBase Model=Qwen2.5-VL-3B, Token Budget=22.2%2025.08 | 73.7 | |
| HiPrune++Base Model=Qwen2.5-VL-3B, Token Budget=22.2%2025.08 | 73.5 | |
| FusedModel size=7B, Evaluation protocol=0-shot2026.02 | 73.28 | |
| UniformModel size=7B, Evaluation protocol=0-shot2026.02 | 71.05 | |
| VRCDdecoding length (L)=384, forward ratio (FR)=0.52026.05 | 69.87 | |
| HiPrune++Base Model=Qwen2.5-VL-3B, Token Budget=11.1%2025.08 | 69.8 | |
| HiPruneBase Model=Qwen2.5-VL-3B, Token Budget=11.1%2025.08 | 69.7 | |
| Entropydecoding length (L)=384, forward ratio (FR)=0.52026.05 | 69.02 | |
| Margindecoding length (L)=384, forward ratio (FR)=0.52026.05 | 68.48 | |
| HyperSegLLM=Phi-2-2.7B2024.11 | 67.9 | |
| IGdecoding length (L)=384, forward ratio (FR)=0.52026.05 | 67.73 | |
| Confidencedecoding length (L)=384, forward ratio (FR)=0.52026.05 | 66.91 | |
| VRCDdecoding length (L)=384, forward ratio (FR)=0.252026.05 | 66.73 | |
| Margindecoding length (L)=384, forward ratio (FR)=0.252026.05 | 66.43 | |
| HoloVRetained Tokens=192, Pruning Ratio=↓ 66.7%2026.04 | 65.4 | |
| LLaVA-1.5-7BRetained Tokens=576, Pruning Ratio=100%2026.04 | 64.7 | |
| LLaVA-1.5LLM=Vicuna-7B2024.11 | 64.3 | |
| HoloVRetained Tokens=128, Pruning Ratio=↓ 77.8%2026.04 | 63.9 | |
| V2DropRetained Tokens=192, Pruning Ratio=↓ 66.7%2026.04 | 63.7 | |
| DeSAPRetained Tokens=192, Pruning Ratio=↓ 66.7%2026.04 | 63.5 | |
| IGdecoding length (L)=384, forward ratio (FR)=0.252026.05 | 63.35 | |
| HoloVRetained Tokens=64, Pruning Ratio=↓ 88.9%2026.04 | 63.3 | |
| PyramidDropRetained Tokens=192, Pruning Ratio=↓ 66.7%2026.04 | 63.3 | |
| FlowCutRetained Tokens=192, Pruning Ratio=↓ 66.7%2026.04 | 63.2 | |
| Confidencedecoding length (L)=384, forward ratio (FR)=0.252026.05 | 63.04 | |
| VisionZipRetained Tokens=192, Pruning Ratio=↓ 66.7%2026.04 | 63 | |
| DeSAPRetained Tokens=128, Pruning Ratio=↓ 77.8%2026.04 | 62.8 | |
| SparseVLMRetained Tokens=192, Pruning Ratio=↓ 66.7%2026.04 | 62.5 | |
| DeSAPRetained Tokens=64, Pruning Ratio=↓ 88.9%2026.04 | 62.3 | |
| FlowCutRetained Tokens=128, Pruning Ratio=↓ 77.8%2026.04 | 62.1 | |
| VisionZipRetained Tokens=128, Pruning Ratio=↓ 77.8%2026.04 | 62 | |
| V2DropRetained Tokens=128, Pruning Ratio=↓ 77.8%2026.04 | 61.8 | |
| PyramidDropRetained Tokens=128, Pruning Ratio=↓ 77.8%2026.04 | 61.6 | |
| VRCDdecoding length (L)=384, forward ratio (FR)=0.1252026.05 | 61.24 | |
| FlowCutRetained Tokens=64, Pruning Ratio=↓ 88.9%2026.04 | 60.8 | |
| Margindecoding length (L)=384, forward ratio (FR)=0.1252026.05 | 60.64 | |
| Qwen-VL-ChatLLM=Qwen-7B2024.11 | 60.6 | |
| VisionZipRetained Tokens=64, Pruning Ratio=↓ 88.9%2026.04 | 60.1 | |
| SparseVLMRetained Tokens=128, Pruning Ratio=↓ 77.8%2026.04 | 60 | |
| LLaVA-PrumergeRetained Tokens=192, Pruning Ratio=↓ 66.7%2026.04 | 59.6 | |
| ShikraLLM=Vicuna-13B2024.11 | 58.8 | |
| PyramidDropRetained Tokens=64, Pruning Ratio=↓ 88.9%2026.04 | 58.8 | |
| LLaVA-PrumergeRetained Tokens=128, Pruning Ratio=↓ 77.8%2026.04 | 58.1 | |
| IGdecoding length (L)=384, forward ratio (FR)=0.1252026.05 | 57.54 | |
| Entropydecoding length (L)=384, forward ratio (FR)=0.252026.05 | 57.48 | |
| Confidencedecoding length (L)=384, forward ratio (FR)=0.1252026.05 | 57.29 | |
| SparseVLMRetained Tokens=64, Pruning Ratio=↓ 88.9%2026.04 | 56.2 | |
| URSAModel Scale=8B, Category=Open-Source Math MLLMs2025.12 | 55.5 | |
| LLaVA-PrumergeRetained Tokens=64, Pruning Ratio=↓ 88.9%2026.04 | 55.3 | |
| V2DropRetained Tokens=64, Pruning Ratio=↓ 88.9%2026.04 | 55.2 | |
| Entropydecoding length (L)=384, forward ratio (FR)=0.1252026.05 | 50.3 | |
| Qwen-VLLLM=Qwen-7B2024.11 | 38.2 | |
| InstructBLIPLLM=Vicuna-7B2024.11 | 36 | |
| FusedModel size=0.5B, Evaluation protocol=0-shot2026.02 | 35.14 | |
| MaD-MixModel size=0.5B, Evaluation protocol=0-shot2026.02 | 34.45 | |
| UniformModel size=0.5B, Evaluation protocol=0-shot2026.02 | 34.36 | |
| AvgModel size=0.5B, Evaluation protocol=0-shot2026.02 | 26.8 |