Visual Question Answering Accuracy on GQA
77.5AccuracyOracle
Evaluation Results
| Method | Links | |
|---|---|---|
| OracleStrategy=Oracle2026.05 | 77.5 | |
| VanillaVisual Token Budget=2880, Token Retention Ratio=100%2026.06 | 64.2 | |
| LLaVA-1.5 + FG-CLIP 2Vision Encoder=FG-CLIP 2, LMM Base Architecture=LLaVA-1.52025.10 | 64 | |
| DeFactoBackbone=Qwen2.5-VL-7B2025.09 | 63.9 | |
| VILA-13BModel Scale=13B, Precision=FP162023.06 | 63.6 | |
| VILA-13B-AWQModel Scale=13B, Quantization=INT4-g1282023.06 | 63.6 | |
| LBR (ours)Backbone=LLaVA-1.5, Parameter Scale=13B, LBR Mitigation=true2026.05 | 63.6 | |
| LLaVA-1.5-13BBackbone=LLaVA-1.5, Parameter Scale=13B, LBR Mitigation=false2026.05 | 63.4 | |
| LLaVA-v1.5-13BModel scale=13B, Number of vision tokens=5762025.08 | 63.3 | |
| VILA-7BModel Scale=7B, Precision=FP162023.06 | 63.1 | |
| VILA-7B-AWQModel Scale=7B, Quantization=INT4-g1282023.06 | 63 | |
| LLaVA Vicuna BaselineMLLM=LLaVA, LLM=Vicuna, Method=Baseline, Activated Params=0.32B2025.09 | 63 | |
| Fourier-LLaVAModel scale=7B, Number of vision tokens=256, Compression ratio=55.6%2025.08 | 62.7 | |
| Fourier-LLaVAModel scale=13B, Number of vision tokens=144, Compression ratio=75.0%2025.08 | 62.7 | |
| LBR (ours)Backbone=LLaVA-1.5, Parameter Scale=7B, LBR Mitigation=true2026.05 | 62.7 | |
| LLaVA-MORELLM=Gemma-2-2B, #TS=1.2M2026.05 | 62.4 | |
| LBR (ours)Backbone=LLaVA-NEXT, Parameter Scale=3B, LBR Mitigation=true2026.05 | 62.4 | |
| LLaVA-1.5 + Meta CLIP 2Vision Encoder=Meta CLIP 2, LMM Base Architecture=LLaVA-1.52025.10 | 62.4 | |
| CompoDistillLLM=Qwen1.5-1.8B, #TS=1.2M2026.05 | 62.2 | |
| DiDi-Merging2026.05 | 62.08 | |
| LLaVA-v1.5-7BModel scale=7B, Number of vision tokens=5762025.08 | 62 | |
| LLaVA-1.5-7BBackbone=LLaVA-1.5, Parameter Scale=7B, LBR Mitigation=false2026.05 | 62 | |
| LLaVA-1.5 + SigLIP 2Vision Encoder=SigLIP 2, LMM Base Architecture=LLaVA-1.52025.10 | 62 | |
| FREE-Merging2026.05 | 61.92 | |
| ImpLLM=Qwen1.5-1.8B, #TS=1.5M2026.05 | 61.9 | |
| VanillaRetained Tokens=576, Backbone=LLaVA-v1.5-7B2026.05 | 61.9 | |
| LLaVA-NEXT-3BBackbone=LLaVA-NEXT, Parameter Scale=3B, LBR Mitigation=false2026.05 | 61.9 | |
| LLaVA-1.5 + CLIPVision Encoder=CLIP, LMM Base Architecture=LLaVA-1.52025.10 | 61.9 | |
| LLaVA-1.5-7BVTR Type=N/A, Tokens Retained=576, Reduction Ratio=100%2026.06 | 61.9 | |
| VanillaToken Budget (K)=576, Token Reduction (%)=100%, Model Backbone=LLaVA-1.5-7B2026.06 | 61.9 | |
| VanillaToken Budget=576, Base Model=LLaVA-1.5-7B2026.06 | 61.9 | |
| MMER2026.05 | 61.85 | |
| Fourier-LLaVAModel scale=7B, Number of vision tokens=144, Compression ratio=75.0%2025.08 | 61.7 | |
| VanillaAverage Token Reduction=0.0%, Backbone=Qwen3-VL-8B2026.06 | 61.7 | |
| MQT-LLaVAModel scale=7B, Number of vision tokens=256, Compression ratio=55.6%2025.08 | 61.6 | |
| Individuals2026.05 | 61.52 | |
| MoE-LLaVALLM=Qwen1.5-1.8B, #TS=2.1M2026.05 | 61.5 | |
| MoE-LLaVA Qwen BaselineMLLM=MoE-LLaVA, LLM=Qwen, Method=Baseline, Activated Params=1.80B2025.09 | 61.5 | |
| CLSEToken Budget=192, Base Model=LLaVA-1.5-7B2026.06 | 61.5 | |
| MQT-LLaVAModel scale=7B, Number of vision tokens=144, Compression ratio=75.0%2025.08 | 61.4 | |
| LLaVA-CKDLLM=Qwen2.5-1.5B, #TS=1.2M2026.05 | 61.2 | |
| HoloV + RESTOREVTR Type=Text-agnostic, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 61 | |
| DivPrune + RESTOREVTR Type=Text-agnostic, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 60.9 | |
| VisPruner + RESTOREVTR Type=Text-agnostic, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 60.9 | |
| VisPruner + RESTOREVTR Type=Text-agnostic, Tokens Retained=128, Reduction Ratio=22.2%2026.06 | 60.9 | |
| CLSEToken Budget=128, Base Model=LLaVA-1.5-7B2026.06 | 60.9 | |
| HoloV + RESTOREVTR Type=Text-agnostic, Tokens Retained=128, Reduction Ratio=22.2%2026.06 | 60.8 | |
| MiniGeminiLLM=Gemma-2B, #TS=2.2M2026.05 | 60.7 | |
| PriorTRVisual Token Budget=320, Token Retention Ratio=11.1%2026.06 | 60.7 | |
| Cloud-OnlyStrategy=Cloud-Only2026.05 | 60.6 | |
| VisionZip + RESTOREVTR Type=Text-agnostic, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 60.6 | |
| DivPrune + RESTOREVTR Type=Text-agnostic, Tokens Retained=128, Reduction Ratio=22.2%2026.06 | 60.6 | |
| Fourier-LLaVAModel scale=7B, Number of vision tokens=64, Compression ratio=88.9%2025.08 | 60.4 | |
| PriorTRToken Budget (K)=192, Token Reduction (%)=↓ 66.7%, Model Backbone=LLaVA-1.5-7B2026.06 | 60.4 | |
| LRCPRetained Tokens=192, Backbone=LLaVA-v1.5-7B2026.05 | 60.3 | |
| MoE-LLaVA StableLM BaselineMLLM=MoE-LLaVA, LLM=StableLM, Method=Baseline, Activated Params=1.60B2025.09 | 60.3 | |
| EAGLEAgent=Qwen, GLM, InternVL2026.05 | 60.3 | |
| ApETRetained Tokens=192, Backbone=LLaVA-v1.5-7B2026.05 | 60.2 | |
| Allign-KDLLM=MobVLMv2-1.7B, #TS=3.6M2026.05 | 60.1 | |
| MQT-LLaVAModel scale=7B, Number of vision tokens=64, Compression ratio=88.9%2025.08 | 60 | |
| INAR-VLStrategy=INAR-VL2026.05 | 59.9 | |
| Mipha Phi-1.5 BaselineMLLM=Mipha, LLM=Phi-1.5, Method=Baseline, Activated Params=1.30B2025.09 | 59.8 | |
| VisionZip + RESTOREVTR Type=Text-agnostic, Tokens Retained=128, Reduction Ratio=22.2%2026.06 | 59.8 | |
| CASTRatio=1%, Sampling Strategy=CAST, Backbone=Qwen2-VL-2B2026.05 | 59.71 | |
| LLaVA Vicuna STSMLLM=LLaVA, LLM=Vicuna, Method=STS, Activated Params=0.27B, FLOPs=82.2%2025.09 | 59.7 | |
| BunnyLLM=Qwen1.5-1.8B, #TS=2.6M2026.05 | 59.6 | |
| LVIDALLM=Llama-3.2-1B, #TS=1.2M2026.05 | 59.6 | |
| VisPrunerToken Budget (K)=192, Token Reduction (%)=↓ 66.7%, Model Backbone=LLaVA-1.5-7B2026.06 | 59.6 | |
| ATP-LLaVAModel scale=7B, Number of vision tokens=144, Compression ratio=75.0%2025.08 | 59.5 | |
| Fourier-LLaVAModel scale=7B, Number of vision tokens=36, Compression ratio=93.75%2025.08 | 59.5 | |
| SparseVLMVTR Type=Text-aware, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 59.5 | |
| ToMeVTR Type=Text-agnostic, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 59.5 | |
| CASTRatio=5%, Sampling Strategy=CAST, Backbone=Qwen2-VL-2B2026.05 | 59.43 | |
| VisPrunerVTR Type=Text-agnostic, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 59.4 | |
| CASTRatio=10%, Sampling Strategy=CAST, Backbone=Qwen2-VL-2B2026.05 | 59.33 | |
| VisionZipRetained Tokens=192, Backbone=LLaVA-v1.5-7B2026.05 | 59.3 | |
| Mipha Phi-1.5 STSMLLM=Mipha, LLM=Phi-1.5, Method=STS, Activated Params=1.08B, FLOPs=83.5%2025.09 | 59.2 | |
| Static routingStrategy=Static routing2026.05 | 59.2 | |
| VisionZipVTR Type=Text-agnostic, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 59.2 | |
| VisPruner + RESTOREVTR Type=Text-agnostic, Tokens Retained=64, Reduction Ratio=11.1%2026.06 | 59.2 | |
| PriorTRToken Budget (K)=128, Token Reduction (%)=↓ 77.8%, Model Backbone=LLaVA-1.5-7B2026.06 | 59.1 | |
| Mipha Phi-2 BaselineMLLM=Mipha, LLM=Phi-2, Method=Baseline, Activated Params=2.70B2025.09 | 59 | |
| Text-OnlyStrategy=Text-Only2026.05 | 59 | |
| DivPrune + RESTOREVTR Type=Text-agnostic, Tokens Retained=64, Reduction Ratio=11.1%2026.06 | 59 | |
| HoloV + RESTOREVTR Type=Text-agnostic, Tokens Retained=64, Reduction Ratio=11.1%2026.06 | 59 | |
| RandomRatio=10%, Sampling Strategy=Random, Backbone=Qwen2-VL-2B2026.05 | 58.98 | |
| ApETRetained Tokens=128, Backbone=LLaVA-v1.5-7B2026.05 | 58.9 | |
| LRCPRetained Tokens=128, Backbone=LLaVA-v1.5-7B2026.05 | 58.9 | |
| DivPruneVTR Type=Text-agnostic, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 58.9 | |
| MQT-LLaVAModel scale=7B, Number of vision tokens=36, Compression ratio=93.75%2025.08 | 58.8 | |
| LLaVA-MoDLLM=Qwen1.5-1.8B, #TS=5M2026.05 | 58.7 | |
| LLaVA-CKDLLM=Qwen2.5-0.5B, #TS=1.2M2026.05 | 58.7 | |
| ToMeVTR Type=Text-agnostic, Tokens Retained=128, Reduction Ratio=22.2%2026.06 | 58.7 | |
| HiREDToken Budget=192, Base Model=LLaVA-1.5-7B2026.06 | 58.7 | |
| RandomRatio=1%, Sampling Strategy=Random, Backbone=Qwen2-VL-2B2026.05 | 58.62 | |
| HoloVVTR Type=Text-agnostic, Tokens Retained=192, Reduction Ratio=33.3%2026.06 | 58.6 | |
| DivPruneVTR Type=Text-agnostic, Tokens Retained=128, Reduction Ratio=22.2%2026.06 | 58.6 | |
| V2DropRetained Tokens=192, Backbone=LLaVA-v1.5-7B2026.05 | 58.5 | |
| VisionZipVTR Type=Text-agnostic, Tokens Retained=128, Reduction Ratio=22.2%2026.06 | 58.5 | |
| VisionZip + RESTOREVTR Type=Text-agnostic, Tokens Retained=64, Reduction Ratio=11.1%2026.06 | 58.5 |