Visual Question Answering on GQA (test)
89.3AccuracyHumans
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Humans2020.04 | 89.3 | 91.2 | 87.4 | — | — | — | — | |
| Kakao*Mode=Ensemble2019.07 | 73.33 | — | — | — | — | — | — | |
| 270Mode=Ensemble2019.07 | 70.23 | — | — | — | — | — | — | |
| NSMMode=Ensemble2019.07 | 67.25 | — | — | — | — | — | — | |
| LLaVA-13BNumber of Parameters=13B2024.05 | 67.1 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=1.2M+1.5M, Res.=384 AnyRes, LLM=Yi-34B2024.05 | 66.6 | — | — | — | — | — | — | |
| LLaVA-NeXTPT+IT=0.5M+0.7M, Res.=336 AnyRes, LLM=Vicuna-13B2024.05 | 65.4 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Vicuna-13B2024.05 | 65.4 | — | — | — | — | — | — | |
| Upper BoundBase Model=LLaVA-NeXT-13B2026.04 | 65.4 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Llama3-8B2024.05 | 65.1 | — | — | — | — | — | — | |
| LLaMA-VIDPT+IT=0.8M+0.7M, Res.=336, LLM=Vicuna-7B2024.05 | 65 | — | — | — | — | — | — | |
| MATAType=Compositional, Agentic type=multi-agent, Setting=General2026.01 | 64.9 | — | — | — | — | — | — | |
| ShareGPT4VPT+IT=1.2M+0.7M, Res.=336, LLM=Vicuna-13B2024.05 | 64.8 | — | — | — | — | — | — | |
| MATAType=Compositional, Agentic type=multi-agent, Setting=Domain-Specific2026.01 | 64.7 | — | — | — | — | — | — | |
| MobileVLM V2PT+IT=1.2M+3.6M, Res.=336, LLM=Vicuna-7B2024.05 | 64.6 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=1.2M+1.5M, Res.=384, LLM=Vicuna-13B2024.05 | 64.6 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Vicuna-7B2024.05 | 64.4 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=1.2M+1.5M, Res.=384 AnyRes, LLM=Vicuna-13B2024.05 | 64.3 | — | — | — | — | — | — | |
| LLaVA-NeXT-7BRetained Tokens=28802025.09 | 64.2 | — | — | — | — | — | — | |
| CLASPBase Model=LLaVA-NeXT-13B, Token Retention Budget=640 Tokens2026.04 | 64.2 | — | — | — | — | — | — | |
| Finetuned mPLUG-OwlNumber of Parameters=7B, Evaluation Protocol=Zero-shot2023.06 | 64 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Llama3-70B LORA2024.05 | 64 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Yi-34B LORA2024.05 | 63.9 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=1.2M+1.5M, Res.=384 AnyRes, LLM=Vicuna-7B2024.05 | 63.9 | — | — | — | — | — | — | |
| InternVL3.5Type=Monolithic, Agentic type=non-agentic, Parameters=8B2026.01 | 63.8 | — | — | — | — | — | — | |
| LLaVA-LLaMA3PT+IT=0.5M+0.6M, Res.=336, LLM=Llama3-8B2024.05 | 63.5 | — | — | — | — | — | — | |
| Mini-GeminiPT+IT=1.2M+1.5M, Res.=336+768, LLM=Vicuna-13B2024.05 | 63.4 | — | — | — | — | — | — | |
| LOVA³-7BLLM=Vicuna-7B2024.05 | 63.4 | — | — | — | — | — | — | |
| LLaVA-v1.5PT+IT=0.5M+0.6M, Res.=336, LLM=Vicuna-13B2024.05 | 63.3 | — | — | — | — | — | — | |
| VILAPT+IT=50M+1M, Res.=336, LLM=Llama-2-13B2024.05 | 63.3 | — | — | — | — | — | — | |
| NSMscene-graph supervision=Krishna et al. (2017)2020.04 | 63.2 | 78.9 | 49.3 | — | — | — | — | |
| CuMoPT+IT=0.5M+0.6M, Res.=336, LLM=Mistral-7B2024.05 | 63.2 | — | — | — | — | — | — | |
| CLASPBase Model=LLaVA-NeXT-13B, Token Retention Budget=320 Tokens2026.04 | 63 | — | — | — | — | — | — | |
| LXRTMode=Ensemble2019.07 | 62.71 | — | — | — | — | — | — | |
| SparseVLMBase Model=LLaVA-NeXT-13B, Token Retention Budget=640 Tokens2026.04 | 62.7 | — | — | — | — | — | — | |
| ChatSearcher2024.10 | 62.5 | — | — | — | — | — | — | |
| Qwen2.5-VLType=Monolithic, Agentic type=non-agentic, Parameters=7B2026.01 | 62.4 | — | — | — | — | — | — | |
| InternVL3Type=Monolithic, Agentic type=non-agentic, Parameters=8B2026.01 | 62.4 | — | — | — | — | — | — | |
| IVM-Enhanced LLaVA-7BNumber of Parameters=14B2024.05 | 62.2 | — | — | — | — | — | — | |
| ZOO-PruneRetained Tokens=6402025.09 | 62.19 | — | — | — | — | — | — | |
| Qwen2-VL-InstructMethod Class=Baseline2025.05 | 62.18 | — | — | — | — | — | — | |
| Qwen2-VL-7B-GRPO-8kCheckpoint Type=Fine-tuned2025.05 | 62.04 | — | — | — | — | — | — | |
| InstructBLIPNumber of Parameters=13B, Evaluation Protocol=Zero-shot2023.06 | 62 | — | — | — | — | — | — | |
| LLaVA-7BNumber of Parameters=7B2024.05 | 62 | — | — | — | — | — | — | |
| LLaVA-1.5LLM=Vicuna-7B2024.05 | 62 | — | — | — | — | — | — | |
| LLaVA-1.5-7BVersion=1.5, Parameters=7B2024.10 | 62 | — | — | — | — | — | — | |
| LLaVA-1.5-7BRetained Tokens=576, Pruning Ratio=0%, Base Model=LLaVA-1.5-7B2025.03 | 61.9 | — | — | — | 100 | — | — | |
| CLASPBase Model=LLaVA-NeXT-13B, Token Retention Budget=160 Tokens2026.04 | 61.8 | — | — | — | — | — | — | |
| DivPruneRetained Tokens=6402025.09 | 61.58 | — | — | — | — | — | — | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Phi2-2.7B2024.05 | 61.5 | — | — | — | — | — | — | |
| InternVL2.5Type=Monolithic, Agentic type=non-agentic, Parameters=8B2026.01 | 61.5 | — | — | — | — | — | — | |
| olmOCR-7B-0225-previewCheckpoint Type=Fine-tuned2025.05 | 61.48 | — | — | — | — | — | — | |
| TSV MergingMethod Class=Merging2025.05 | 61.4 | — | — | — | — | — | — | |
| Iso-CMethod Class=Merging2025.05 | 61.34 | — | — | — | — | — | — | |
| Mixture TrainingProtocol=full fine-tuning2025.05 | 61.33 | — | — | — | — | — | — | |
| TinyLLaVAPT+IT=0.5M+0.6M, Res.=384, LLM=Phi2-2.7B2024.05 | 61.3 | — | — | — | — | — | — | |
| VisionZipRetained Tokens=6402025.09 | 61.3 | — | — | — | — | — | — | |
| OptMergeMethod Class=Merging2025.05 | 61.29 | — | — | — | — | — | — | |
| GRNMode=Ensemble2019.07 | 61.22 | — | — | — | — | — | — | |
| IPCV+FastVBackbone=InternVL3-38B, ViT token retention rate=40%, LLM token retention rate=40%2025.12 | 61.2 | — | — | — | — | 89.9 | 0.644 | |
| TwigVLMRetained Tokens=192, Pruning Ratio=66.7%, Base Model=LLaVA-1.5-7B2025.03 | 61.2 | — | — | — | 99.2 | — | — | |
| TwigVLM++Retained Tokens=192, Pruning Ratio=66.7%, Base Model=LLaVA-1.5-7B2025.03 | 61.2 | — | — | — | 99.6 | — | — | |
| TA w/ DAREMethod Class=Merging2025.05 | 61.15 | — | — | — | — | — | — | |
| MobileVLM V2PT+IT=1.2M+3.6M, Res.=336, LLM=ML-2.7B2024.05 | 61.1 | — | — | — | — | — | — | |
| UniTokLLM=LLaMA-2-7B, PT Data Size=70M2025.05 | 61.1 | — | — | — | — | — | — | |
| MSMMode=Ensemble2019.07 | 61.09 | — | — | — | — | — | — | |
| ZOO-PruneRetained Tokens=3202025.09 | 60.97 | — | — | — | — | — | — | |
| Qwen2-VL-7B-PokemonCheckpoint Type=Fine-tuned2025.05 | 60.96 | — | — | — | — | — | — | |
| Task ArithmeticMethod Class=Merging2025.05 | 60.95 | — | — | — | — | — | — | |
| DREAMMode=Ensemble2019.07 | 60.93 | — | — | — | — | — | — | |
| Individual VQAProtocol=full fine-tuning2025.05 | 60.91 | — | — | — | — | — | — | |
| SparseVLMBase Model=LLaVA-NeXT-13B, Token Retention Budget=320 Tokens2026.04 | 60.9 | — | — | — | — | — | — | |
| SK T-Brain*Mode=Ensemble2019.07 | 60.87 | — | — | — | — | — | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B, Token Retention Ratio=Full Tokens2026.06 | 60.84 | — | — | — | 100 | — | — | |
| VanillaBackbone=InternVL3-38B, ViT token retention rate=100%, LLM token retention rate=100%2025.12 | 60.8 | — | — | — | — | 100 | 1 | |
| DART (ViT)Backbone=InternVL3-38B, ViT token retention rate=40%, LLM token retention rate=40%2025.12 | 60.8 | — | — | — | — | 88.1 | 0.639 | |
| TwigVLM++Retained Tokens=128, Pruning Ratio=77.8%, Base Model=LLaVA-1.5-7B2025.03 | 60.8 | — | — | — | 99.2 | — | — | |
| PKUMode=Ensemble2019.07 | 60.79 | — | — | — | — | — | — | |
| TIES w/ DAREMethod Class=Merging2025.05 | 60.78 | — | — | — | — | — | — | |
| TwigVLMRetained Tokens=128, Pruning Ratio=77.8%, Base Model=LLaVA-1.5-7B2025.03 | 60.6 | — | — | — | 98.7 | — | — | |
| TIES MergingMethod Class=Merging2025.05 | 60.54 | — | — | — | — | — | — | |
| STSBackbone=Qwen2.5-VL-7B, Token Retention Ratio=20%2026.06 | 60.31 | — | — | — | 97.1 | — | — | |
| SparseVLMRetained Tokens=6402025.09 | 60.3 | — | — | — | — | — | — | |
| ToMeBackbone=InternVL3-38B, ViT token retention rate=40%, LLM token retention rate=40%2025.12 | 60.2 | — | — | — | — | 82.7 | 0.617 | |
| WUDI MergingMethod Class=Merging2025.05 | 60.11 | — | — | — | — | — | — | |
| VisionZip‡Retained Tokens=192, Pruning Ratio=66.7%, Base Model=LLaVA-1.5-7B2025.03 | 60.1 | — | — | — | 98.3 | — | — | |
| DivPruneBackbone=Qwen2.5-VL-7B, Token Retention Ratio=20%2026.06 | 60.05 | — | — | — | 96 | — | — | |
| Finetuned mPLUG-Owl (Pseudo Instruction)Number of Parameters=7B, Evaluation Protocol=Zero-shot, Training strategy=Finetuned on pseudo instruction data2023.06 | 60 | — | — | — | — | — | — | |
| Task Arithmeticunsupervised=false2025.03 | 59.98 | — | — | — | — | — | — | |
| AdaMMSunsupervised=true2025.03 | 59.96 | — | — | — | — | — | — | |
| MusanMode=Ensemble2019.07 | 59.93 | — | — | — | — | — | — | |
| MetaGPTunsupervised=true2025.03 | 59.93 | — | — | — | — | — | — | |
| ZOO-PruneRetained Tokens=1602025.09 | 59.93 | — | — | — | — | — | — | |
| ToFuBackbone=InternVL3-38B, ViT token retention rate=40%, LLM token retention rate=40%2025.12 | 59.8 | — | — | — | — | 83.4 | 0.619 | |
| TwigVLM++Retained Tokens=64, Pruning Ratio=88.9%, Base Model=LLaVA-1.5-7B2025.03 | 59.7 | — | — | — | 97.7 | — | — | |
| DivPruneRetained Tokens=3202025.09 | 59.63 | — | — | — | — | — | — | |
| CogVLM(base)Model Type=Original2025.03 | 59.43 | — | — | — | — | — | — | |
| SparseVLMBase Model=LLaVA-NeXT-13B, Token Retention Budget=160 Tokens2026.04 | 59.4 | — | — | — | — | — | — | |
| VisionZipRetained Tokens=3202025.09 | 59.3 | — | — | — | — | — | — | |
| Qwen-VL2024.10 | 59.3 | — | — | — | — | — | — |