Visual Question Answering on GQA (GQA Score and Average)
63.4GQA ScoreInternVL 2.5 8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVL 2.5 8BFLOPs=12.575, Ratio(%)=100.0, Token budget=Full2026.03 | 63.4 | 100 | |
| ShareGPT4VLLM=Vicuna-7B, #Samples=1.9M2024.12 | 63.3 | 70.8 | |
| Full-DataSampling Ratio=100% (665K), Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 63 | — | |
| BF16Precision=BF162026.06 | 62.95 | — | |
| TokenFlow-XLType=Und. & Gen., # Params=13B2025.09 | 62.7 | — | |
| MobileVLM V2 7BLLM=Vicuna-7B, #Samples=3.6M2024.12 | 62.6 | 72.2 | |
| ASAPFLOPs=3.969, Ratio(%)=31.56, Token budget=510 tokens2026.03 | 62.31 | 98.76 | |
| LLaVA-OneVision 7B, DenseBackbone=LLaVA-OneVision 7B, Sparsity=Dense2026.03 | 62.25 | — | |
| LLaVA-v1.5Type=Und. Only, # Params=7B2025.09 | 62 | — | |
| LLaVA-1.5LLM=Vicuna-7B, #Samples=1.2M2024.12 | 62 | 68.8 | |
| FastVFLOPs=3.969, Ratio(%)=31.56, Token budget=510 tokens2026.03 | 61.44 | 97.17 | |
| XMASSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 61.2 | — | |
| MobileVLM-V2Type=Und. Only, # Params=2.7B2025.09 | 61.1 | — | |
| MoE-LLaVA-2.7B×4LLM=Phi-2.7B, #Samples=2.2M2024.12 | 61.1 | 66.7 | |
| MobileVLM V2 3BLLM=MobileLLaMA 2.7B, #Samples=3.6M2024.12 | 61.1 | 68.1 | |
| CoIDOSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 61 | — | |
| VILA-UType=Und. & Gen., # Params=7B2025.09 | 60.8 | — | |
| ICONSSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 60.7 | — | |
| Qwen2.5-VL 7B, DenseBackbone=Qwen2.5-VL 7B, Sparsity=Dense2026.03 | 60.49 | — | |
| MoE-LLaVA-1.6B×4LLM=StableLM-1.6B, #Samples=2.2M2024.12 | 60.4 | 63.3 | |
| Emu3-ChatType=Und. Only, # Params=8B2025.09 | 60.3 | — | |
| OFASampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 60.3 | — | |
| COINCIDESampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 60.2 | — | |
| MODEQuantization=W32026.06 | 60.15 | — | |
| MobileVLM V2 1.7BSubset=Long, #Samples=3.6M, Align-KD=true2024.12 | 60.1 | 65.1 | |
| ASAPFLOPs=2.663, Ratio(%)=21.17, Token budget=340 tokens2026.03 | 60 | 95.69 | |
| TypiClustSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 59.8 | — | |
| MobileVLM-V2Type=Und. Only, # Params=1.4B2025.09 | 59.3 | — | |
| Janus-Pro-1BType=Und. & Gen., # Params=1.5B2025.09 | 59.3 | — | |
| VEQ-MAQuantization=W32026.06 | 59.21 | — | |
| D-DitType=Und. & Gen., # Params=2.0B2025.09 | 59.2 | — | |
| FastVFLOPs=2.663, Ratio(%)=21.17, Token budget=340 tokens2026.03 | 59.17 | 95.02 | |
| JanusType=Und. & Gen., # Params=1.5B2025.09 | 59.1 | — | |
| MobileVLMType=Und. Only, # Params=2.7B2025.09 | 59 | — | |
| MobileVLM 3BLLM=MobileLLaMA 2.7B, #Samples=3.6M2024.12 | 59 | 62.8 | |
| MobileVLM V2 1.7BSubset=Long, #Samples=3.6M, Align-KD=false2024.12 | 59 | 63.7 | |
| MobileVLM V2 1.7BSubset=Short, #Samples=3.6M, Align-KD=true2024.12 | 58.9 | 64.4 | |
| SparseGPTBackbone=LLaVA-OneVision 7B, Sparsity=60% Sparsity2026.03 | 58.56 | — | |
| MODEQuantization=W22026.06 | 58.4 | — | |
| OursType=Und. & Gen., # Params=1.5B2025.09 | 58.2 | — | |
| ATV-PruningBackbone=LLaVA-OneVision 7B, Sparsity=60% Sparsity2026.03 | 58.05 | — | |
| Show-o-512Type=Und. & Gen., # Params=1.3B2025.09 | 58 | — | |
| PreSelSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 57.9 | — | |
| IFDSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 57.8 | — | |
| TAMPBackbone=LLaVA-OneVision 7B, Sparsity=60% Sparsity2026.03 | 57.64 | — | |
| FUDOKIType=Und. & Gen., # Params=1.5B2025.09 | 57.6 | — | |
| ATV-PruningBackbone=Qwen2.5-VL 7B, Sparsity=60% Sparsity2026.03 | 57.58 | — | |
| Qwen-VL-ChatType=Und. Only, # Params=7B2025.09 | 57.5 | — | |
| MC-MoEQuantization=W32026.06 | 57.38 | — | |
| CLIP-ScoreSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 56.9 | — | |
| LLaVA-Phi-1.5Type=Und. Only, # Params=1.3B2025.09 | 56.5 | — | |
| Self-FilterSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 56.3 | — | |
| MobileVLMType=Und. Only, # Params=1.4B2025.09 | 56.1 | — | |
| MobileVLM 1.7BLLM=MobileLLaMA 1.4B, #Samples=3.6M2024.12 | 56.1 | 58.7 | |
| VEQ-MAQuantization=W22026.06 | 56.04 | — | |
| GPTQQuantization=W32026.06 | 56.01 | — | |
| TAMPBackbone=Qwen2.5-VL 7B, Sparsity=60% Sparsity2026.03 | 55.76 | — | |
| MC-MoEQuantization=W22026.06 | 55.63 | — | |
| RandomSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 55.1 | — | |
| MobileVLM V2 1.7BSubset=Short, #Samples=3.6M, Align-KD=false2024.12 | 55.1 | 62.4 | |
| WandaBackbone=LLaVA-OneVision 7B, Sparsity=60% Sparsity2026.03 | 54.86 | — | |
| EL2NSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 54.7 | — | |
| SparseGPTBackbone=Qwen2.5-VL 7B, Sparsity=60% Sparsity2026.03 | 54.26 | — | |
| WandaBackbone=Qwen2.5-VL 7B, Sparsity=60% Sparsity2026.03 | 51.3 | — | |
| MAGICDataset=Vision-Flan-186K, #Data=∼37k2026.05 | 50.7 | — | |
| InstructBLIPType=Und. Only, # Params=13B2025.09 | 49.5 | — | |
| InstructBLIPType=Und. Only, # Params=7B2025.09 | 49.2 | — | |
| Full-FinetuneDataset=Vision-Flan-186K, #Data=186k2026.05 | 49.2 | — | |
| Show-o-256Type=Und. & Gen., # Params=1.3B2025.09 | 48.7 | — | |
| LaVITType=Und. & Gen., # Params=7B2025.09 | 46.8 | — | |
| RandomDataset=Vision-Flan-186K, #Data=∼37k2026.05 | 45.8 | — | |
| LWMType=Und. & Gen., # Params=7B2025.09 | 44.8 | — | |
| IDEFICSType=Und. Only, # Params=8B2025.09 | 38.4 | — | |
| MiniGPT-4LLM=Vicuna-7B, #Samples=5.0M2024.12 | 32.2 | — | |
| GPTQQuantization=W22026.06 | 30.3 | — |