Visual Question Answering on A-OKVQA
92.68AccCont-Squeeze (128 -> 1)
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Cont-Squeeze (128 -> 1)Optimization steps=+0 steps2026.02 | 92.68 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Cont-Squeeze (128 -> 1)Optimization steps=+200 steps2026.02 | 92.44 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Cont-Squeeze (128 -> 1)Optimization steps=+700 steps2026.02 | 92.37 | — | — | — | — | — | — | — | — | — | — | — | — | |
| In-Squeeze (128 ... -> 1)Optimization mode=Standard2026.02 | 92 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Direct Fine-tuning (r=1)Optimization steps=+0 steps2026.02 | 91.94 | — | — | — | — | — | — | — | — | — | — | — | — | |
| In-Squeeze (128 ... -> 1)Optimization mode=Min steps2026.02 | 91.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Direct Fine-tuning (r=1)Optimization steps=+200 steps2026.02 | 91.33 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Direct Fine-tuning (r=1)Optimization steps=+700 steps2026.02 | 91.27 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaseBase model=InternVL3.5-8B2026.07 | 90 | — | — | — | — | — | — | — | — | — | 4.24 | 4.14 | — | |
| ReShiftBase model=InternVL3.5-8B2026.07 | 90 | — | — | — | — | — | — | — | — | 0.97 | 4.31 | 4.09 | 2 | |
| Gemma 3 4B ITZero-shot=true2026.02 | 89 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OpenVLThinker-7BTraining=Base2026.05 | 88.73 | — | — | — | — | — | — | — | — | — | — | — | — | |
| WeThink-7BTraining=Base2026.05 | 88.65 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8BTraining=Base2026.05 | 87.86 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BTraining=Staged2026.05 | 87.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ALEAHallu2025.12 | 87.24 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaModels=InternVL 2.5 8B, Average Tokens=100%2026.02 | 87.2 | — | — | — | — | 100 | — | — | — | — | — | — | — | |
| IVC-PruneModels=InternVL 2.5 8B, Average Tokens=50%2026.02 | 87.2 | — | — | — | — | 100 | — | — | — | — | — | — | — | |
| ReShiftBase model=Qwen2.5-VL-7B2026.07 | 87 | — | — | — | — | — | — | — | — | 0.97 | 4.03 | 4.01 | 0 | |
| PDropModels=InternVL 2.5 8B, Average Tokens=56%2026.02 | 86.9 | — | — | — | — | 99.1 | — | — | — | — | — | — | — | |
| VanillaModels=DeepSeek-VL2 Small-16B, Average Tokens=100%2026.02 | 86.9 | — | — | — | — | 100 | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BTraining=Base2026.05 | 86.81 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GThinker-7BTraining=Base2026.05 | 86.72 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OneThinker-8BTraining=Base2026.05 | 86.72 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IVC-PruneModels=Qwen2.5-VL 7B, Average Tokens=50%2026.02 | 86.7 | — | — | — | — | 100.2 | — | — | — | — | — | — | — | |
| PDropModels=DeepSeek-VL2 Small-16B, Average Tokens=57%2026.02 | 86.7 | — | — | — | — | 100 | — | — | — | — | — | — | — | |
| FastVModels=InternVL 2.5 8B, Average Tokens=53%2026.02 | 86.6 | — | — | — | — | 98.6 | — | — | — | — | — | — | — | |
| IVC-PruneModels=DeepSeek-VL2 Small-16B, Average Tokens=52%2026.02 | 86.6 | — | — | — | — | 100.1 | — | — | — | — | — | — | — | |
| VanillaModels=Qwen2.5-VL 7B, Average Tokens=100%2026.02 | 86.5 | — | — | — | — | 100 | — | — | — | — | — | — | — | |
| FastVModels=Qwen2.5-VL 7B, Average Tokens=54%2026.02 | 86.4 | — | — | — | — | 98.2 | — | — | — | — | — | — | — | |
| Qwen3-VL-8BTraining=Staged2026.05 | 86.29 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Nullu2025.12 | 86 | — | — | — | — | — | — | — | — | — | — | — | — | |
| D.P.Data Source Type=Real Data Reweighting, Sampling Strategy Category=Ours, Compute Budget=280K datapoints2026.03 | 86 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaseBase model=Qwen2.5-VL-7B2026.07 | 86 | — | — | — | — | — | — | — | — | — | 4.32 | 4.03 | — | |
| FastVModels=DeepSeek-VL2 Small-16B, Average Tokens=54%2026.02 | 85.9 | — | — | — | — | 98.8 | — | — | — | — | — | — | — | |
| OracleData Source Type=Real Data Reweighting, Sampling Strategy Category=Gold, Compute Budget=280K datapoints2026.03 | 85.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PDropModels=Qwen2.5-VL 7B, Average Tokens=61%2026.02 | 85.6 | — | — | — | — | 98 | — | — | — | — | — | — | — | |
| MMR1-7BTraining=Base2026.05 | 85.07 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RewriteBase model=InternVL3.5-8B2026.07 | 85 | — | — | — | — | — | — | — | — | 0.87 | 3.07 | 3.49 | 9 | |
| ICONSData Source Type=Real Data Reweighting, Sampling Strategy Category=Baseline, Compute Budget=280K datapoints2026.03 | 84.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OPERADecoding Strategy=OPERA2025.12 | 84.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Beam SearchDecoding Strategy=Beam Search2025.12 | 84.12 | — | — | — | — | — | — | — | — | — | — | — | — | |
| UniformData Source Type=Real Data Reweighting, Sampling Strategy Category=Baseline, Compute Budget=280K datapoints2026.03 | 83.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| R1-OneVision-RL-7BTraining=Base2026.05 | 83.58 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BadTokenBase model=InternVL3.5-8B2026.07 | 83 | — | — | — | — | — | — | — | — | 0.9 | 2.94 | 3.04 | 13 | |
| DyCo-RLBackbone=Qwen2.5-VL-3B, Zero-shot=true2026.06 | 82.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO BaselineBackbone=Qwen2.5-VL-3B, Zero-shot=true2026.06 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BadTokenBase model=Qwen2.5-VL-7B2026.07 | 82 | — | — | — | — | — | — | — | — | 0.92 | 3.72 | 2.89 | 9 | |
| H-GIVRModel=Qwen2.5vl:7b, Prompt=Ours2026.02 | 81.83 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Full KV CachingModel=LLaVA1.5-13B2026.03 | 81.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| AttentionPackModel=LLaVA1.5-13B, Rkv=64, Rvv=64, Cache Memory Reduction=5.17×, Average Throughput Change=+43%2026.03 | 81.25 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VCDDecoding Strategy=Visual Contrastive Decoding2025.12 | 81.24 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MinicacheModel=LLaVA1.5-13B, Cache Memory Reduction=4.35×, Average Throughput Change=+38%2026.03 | 81.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVModel=LLaVA1.5-13B, k=5, e=50%, Cache Memory Reduction=1×, Average Throughput Change=+20%2026.03 | 81.17 | — | — | — | — | — | — | — | — | — | — | — | — | |
| InstructBLIPParadigm=Supervised2023.10 | 81 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Regular2025.12 | 80.88 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MMBoundaryBase Model=Qwen2VL 7B2025.05 | 80.6 | 30.5 | 34.8 | 66.4 | — | — | — | — | — | — | — | — | — | |
| AttentionPackModel=LLaVA1.5-13B, Rkv=32, Rvv=32, Cache Memory Reduction=7.30×, Average Throughput Change=+50%2026.03 | 80 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RewriteBase model=Qwen2.5-VL-7B2026.07 | 80 | — | — | — | — | — | — | — | — | 0.89 | 3.39 | 3.3 | 4 | |
| RCEBase Model=Qwen2VL 7B2025.05 | 79.3 | 37.2 | 42.7 | 64.7 | — | — | — | — | — | — | — | — | — | |
| PDropModels=LLaVA-v1.5 7B, Average Tokens=47%2026.02 | 79.2 | — | — | — | — | 99.8 | — | — | — | — | — | — | — | |
| H2OModel=LLaVA1.5-13B, e=50%, Cache Memory Reduction=2×, Average Throughput Change=+24%2026.03 | 79.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVModels=LLaVA-v1.5 7B, Average Tokens=30%2026.02 | 79 | — | — | — | — | 99.7 | — | — | — | — | — | — | — | |
| BadVisionBase model=InternVL3.5-8B2026.07 | 79 | — | — | — | — | — | — | — | — | 0.44 | 3.53 | 2.75 | 6 | |
| VanillaModels=LLaVA-v1.5 7B, Average Tokens=100%2026.02 | 78.9 | — | — | — | — | 100 | — | — | — | — | — | — | — | |
| IVC-PruneModels=LLaVA-v1.5 7B, Average Tokens=28%2026.02 | 78.8 | — | — | — | — | 101.3 | — | — | — | — | — | — | — | |
| ScissorHandsModel=LLaVA1.5-13B, e=50%, Cache Memory Reduction=2×, Average Throughput Change=+24%2026.03 | 77.48 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Conf-CSRBase Model=Qwen2VL 7B2025.05 | 77.4 | 41.3 | 46.3 | 62.3 | — | — | — | — | — | — | — | — | — | |
| BadVisionBase model=Qwen2.5-VL-7B2026.07 | 77 | — | — | — | — | — | — | — | — | 0.37 | 2.67 | 2.76 | 7 | |
| AttentionPackModel=LLaVA1.5-7B, Rkv=64, Rvv=64, Cache Memory Reduction=5.09×, Average Throughput Change=+54%2026.03 | 76.88 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVModel=LLaVA1.5-7B, k=5, e=50%, Cache Memory Reduction=1×, Average Throughput Change=+29%2026.03 | 76.72 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Full KV CachingModel=LLaVA1.5-7B2026.03 | 76.64 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MinicacheModel=LLaVA1.5-7B, Cache Memory Reduction=4.51×, Average Throughput Change=+44%2026.03 | 76.54 | — | — | — | — | — | — | — | — | — | — | — | — | |
| H-GIVRModel=Llama3.2-vision:11b, Prompt=Ours2026.02 | 76.42 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ProphetParadigm=Supervised, trained_on_dataset=true2023.10 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| H2OModel=LLaVA1.5-7B, e=50%, Cache Memory Reduction=2×, Average Throughput Change=+32%2026.03 | 76.36 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DRLBase Model=Qwen2VL 7B2025.05 | 76.2 | 40.6 | 44.2 | 62.8 | — | — | — | — | — | — | — | — | — | |
| AttentionPackModel=LLaVA1.5-7B, Rkv=32, Rvv=32, Cache Memory Reduction=7.24×, Average Throughput Change=+65%2026.03 | 75.91 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Full KV CachingModel=QwenVL-Chat-7B2026.03 | 75.67 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ViCorParadigm=In-Context Learning, Backbone=GPT-42023.10 | 75.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ScissorHandsModel=LLaVA1.5-7B, e=50%, Cache Memory Reduction=2×, Average Throughput Change=+32%2026.03 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| AttentionPackModel=QwenVL-Chat-7B, Rkv=64, Rvv=64, Cache Memory Reduction=2.77×, Average Throughput Change=+61%2026.03 | 75.33 | — | — | — | — | — | — | — | — | — | — | — | — | |
| AttentionPackModel=QwenVL-Chat-7B, Rkv=32, Rvv=32, Cache Memory Reduction=4.02×, Average Throughput Change=+74%2026.03 | 75.17 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVModel=QwenVL-Chat-7B, k=2, e=50%, Cache Memory Reduction=1×, Average Throughput Change=+45%2026.03 | 75.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| H2OModel=QwenVL-Chat-7B, e=50%, Cache Memory Reduction=2×, Average Throughput Change=+51%2026.03 | 74.98 | — | — | — | — | — | — | — | — | — | — | — | — | |
| D.P.Data Source Type=Synthetic Data Selection, Sampling Strategy Category=Ours, Compute Budget=280K datapoints2026.03 | 74.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| AssistGPTParadigm=In-Context Learning, Backbone=GPT-42023.10 | 74.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ScissorHandsModel=QwenVL-Chat-7B, e=50%, Cache Memory Reduction=2×, Average Throughput Change=+51%2026.03 | 74.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ICONSData Source Type=Synthetic Data Selection, Sampling Strategy Category=Baseline, Compute Budget=280K datapoints2026.03 | 74.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-1.5-13BModality=I,T, Model Scale=13B2024.05 | 73.54 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PromptCapParadigm=Supervised, trained_on_dataset=true2023.10 | 73.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SaySelfBase Model=Qwen2VL 7B2025.05 | 72.7 | 32.4 | 38.1 | 63.2 | — | — | — | — | — | — | — | — | — | |
| Active-ProModel=Qwen2.5vl:7b, Prompt=Active-Pro2026.02 | 71.26 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-1.5-7BModality=I,T, Model Scale=7B2024.05 | 70.92 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ViCorParadigm=In-Context Learning, Backbone=GPT-3.52023.10 | 70.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Uni-MoE w/ MoE-Task4 (8E) + Aux lossModality=I,T,S,V, Training Task Strategy=MoE-Task4, Number of Experts=8, Auxiliary Balancing Loss=true2024.05 | 70.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4Paradigm=In-Context Learning, Reasoning=Chain-of-Thought2023.10 | 70.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Uni-MoE w/ MoE-Task4 (4E)Modality=I,T,S,V, Training Task Strategy=MoE-Task4, Number of Experts=42024.05 | 70.22 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Uni-MoE w/ MoE-Task4 (8E)Modality=I,T,S,V, Training Task Strategy=MoE-Task4, Number of Experts=82024.05 | 70 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Auto-CoTModel=Qwen2.5vl:7b, Prompt=Auto-CoT2026.02 | 69.78 | — | — | — | — | — | — | — | — | — | — | — | — |