Visual Question Answering on VQAv2
97.4AccuracyOracle
Evaluation Results
| Method | Links | |
|---|---|---|
| OracleStrategy=Oracle2026.05 | 97.4 | |
| SEA-PRIMELLM=Llama3-8B, Res.=3842024.08 | 83.1 | |
| MM1Vision Encoder=CLIP-L, Resolution=1792, LLM=7B, Data=1B2025.01 | 82.8 | |
| PIIP-LLaVAVision Encoder=ConvNeXt-L, CLIP-L, Resolution=1024/336, LLM=Vicuna-13B, Data=1.2M2025.01 | 82.5 | |
| LLAVA-HRVision Encoder=ConvNeXt-L, CLIP-L, Resolution=1024/448, LLM=Vicuna-13B, Data=1.2M2025.01 | 82.3 | |
| PIIP-LLaVAVision Encoder=ConvNeXt-L, CLIP-L, Resolution=1024/336, LLM=Vicuna-7B, Data=2.7M2025.01 | 82.3 | |
| CogVLM# Param=17B2024.08 | 82.3 | |
| MM1-7B-ChatBackbone=MM1-7B, #Data=>2B2024.06 | 82.3 | |
| mPLUG-Owl3# Param=8B2024.08 | 82.1 | |
| MG-LLAVALLM=Yi1.5-34B, Parameters=34.4B, Resolution=336(768)2024.06 | 82 | |
| LLaVA-HRVision Encoder=ConvNeXt-L, CLIP-L, Resolution=1024/448, LLM=Vicuna-7B, Data=1.2M2025.01 | 81.9 | |
| SEA-PRIMELLM=Vicuna-13B, Res.=3842024.08 | 81.9 | |
| EVLM-Chat# Param=32B2024.08 | 81.9 | |
| FP16Backbone=Qwen3-VL-8B2026.06 | 81.84 | |
| PIIP-LLaVAVision Encoder=ConvNeXt-B, CLIP-L, Resolution=1024/336, LLM=Vicuna-13B, Data=1.2M2025.01 | 81.8 | |
| LLaVA-NeXTVision Encoder=CLIP-L, Resolution=1344, LLM=Vicuna-7B, Data=1.6M2025.01 | 81.8 | |
| LLAVA-Next# Param=8B2024.08 | 81.8 | |
| LLaVA-NeXTBackbone=Vicuna-7B, #Data=1.2M2024.06 | 81.8 | |
| LLaVA-UHDLLM=Vicuna-13B, Parameters=13.4B, Resolution=672x10082024.06 | 81.7 | |
| PIIP-LLaVAVision Encoder=ConvNeXt-L, CLIP-L, Resolution=1024/336, LLM=Vicuna-7B, Data=1.2M2025.01 | 81.5 | |
| SEA-PRIMELLM=Vicuna-7B, Res.=3842024.08 | 81.4 | |
| VisionLLM v2-chatBackbone=Vicuna-7B, #Data=22M2024.06 | 81.4 | |
| Flamingo# Param=80B2024.08 | 81.3 | |
| Mipha-3BLM=Phi-2-2.7B, Res.=3842024.12 | 81.3 | |
| MG-LLAVALLM=Vicuna-13B, Parameters=13.6B, Resolution=336(768)2024.06 | 81.2 | |
| PIIP-LLaVAVision Encoder=ConvNeXt-B, CLIP-L, Resolution=1024/336, LLM=Vicuna-7B, Data=1.2M2025.01 | 81.1 | |
| SharQBackbone=Qwen3-VL-8B2026.06 | 81.1 | |
| SEA-PRIMELLM=Gemma-2B, Res.=3842024.08 | 81 | |
| Dragonfly (Ours)Backbone=Llama3-8B, #Data=2.9M2024.06 | 81 | |
| S2-WrapperLLM=Vicuna-13B, Res.=10082024.08 | 80.9 | |
| VILA-1.5# Param=8B2024.08 | 80.9 | |
| VILA-13BLLM=Llama-2-13B, Resolution=336, Pre-training samples=50M, Instruction tuning samples=1M2023.12 | 80.8 | |
| Idefics2# Param=8B2024.08 | 80.8 | |
| DualPDBackbone=Qwen-2.5-VL-7B, Decoding Strategy=DualPD2026.01 | 80.78 | |
| SEA-PRIMELLM=Phi3-3.8B, Res.=3842024.08 | 80.7 | |
| MG-LLAVALLM=LLaMA3-8B, Parameters=8.2B, Resolution=336(768)2024.06 | 80.7 | |
| SPHINX-2kBackbone=Llama2-7B, #Data=1B2024.06 | 80.7 | |
| ShareGPT4VBackbone=Vicuna-7B, Resolution=336, Visual Input Type=Continuous2024.12 | 80.6 | |
| ShareGPT4VLLM=Vicuna-7B, Res.=3362024.08 | 80.6 | |
| VILA-13B + ShareGPT4VLLM=Llama-2-13B, Resolution=336, Pre-training samples=50M, Instruction tuning samples=1M2023.12 | 80.6 | |
| ShareGPT4V-7BBackbone=Vicuna-7B, #Data=1.8M2024.06 | 80.6 | |
| Cloud-OnlyStrategy=Cloud-Only2026.05 | 80.6 | |
| OlympusLM=Phi-2-2.7B, Res.=3842024.12 | 80.5 | |
| NVFP4Backbone=Qwen3-VL-8B2026.06 | 80.47 | |
| SlimeVision Encoder=CLIP-L, Resolution=2016, LLM=Vicuna-7B, Data=2M2025.01 | 80.3 | |
| MonkeyBackbone=Qwen-7B, #Data=1B2024.06 | 80.3 | |
| MG-LLaVAVision Encoder=ConvNeXt-L, CLIP-L, RAMPlus, OWL-ViTv2-L, Resolution=768/336, LLM=Vicuna-7B, Data=2.5M2025.01 | 80.2 | |
| SPHINX-1kLLM=Vicuna-7B, Parameters=10B, Resolution=4482024.06 | 80.2 | |
| MG-LLAVALLM=Vicuna-7B, Parameters=7.4B, Resolution=336(768)2024.06 | 80.2 | |
| HiMAPModel=LLaVA-13B, TFLOPS=1.36, FLOPs Ratio=23%2025.03 | 80.2 | |
| LLaVA-1.5-13B + MMFuserLLM=Vicuna-13B, Training annotations observed=true2024.10 | 80.1 | |
| LLaVA-1.5Vision Encoder=CLIP-L, Resolution=336, LLM=Vicuna-13B, Data=1.2M2025.01 | 80 | |
| LLaVA-1.5Backbone=Vicuna-13B, Resolution=336, Visual Input Type=Continuous2024.12 | 80 | |
| Show-oBackbone=Phi-1.5-1.3B, Resolution=256, Visual Input Type=Discrete2024.12 | 80 | |
| LLaMA-VIDLLM=Vicuna-13B, Res.=3362024.08 | 80 | |
| LLaVA-1.5*LLM=Vicuna-13B, Res.=3362024.08 | 80 | |
| AlignGPTLLM=Vicuna-13B, Res.=3362024.08 | 80 | |
| LLaVA-1.5LLM=Vicuna-1.5-13B, Resolution=336, Pre-training samples=0.6M, Instruction tuning samples=0.7M2023.12 | 80 | |
| LLaVA-1.5LLM=Vicuna-13B, Parameters=13.4B, Resolution=3362024.06 | 80 | |
| LLaVA-1.5-13BLLM=Vicuna-13B, Training annotations observed=true2024.10 | 80 | |
| Full-FinetuneTarget Model=LLaVA-1.5-13B, Data Budget=100%2026.05 | 80 | |
| LLaVA-1.5-13BBackbone=LLaVA-1.5-13B, Retained Tokens=576 (Upper Bound)2026.06 | 80 | |
| TinyLLaVALLM=Phi-2.7B, Res.=3842024.08 | 79.9 | |
| VILA-7BLLM=Llama-2-7B, Resolution=336, Pre-training samples=50M, Instruction tuning samples=1M2023.12 | 79.9 | |
| VILABackbone=Llama2-7B, #Data=61M2024.06 | 79.9 | |
| MoE-LLaVA-3.6BLM=Phi-2-2.7B, Res.=3842024.12 | 79.9 | |
| TinyLLaVALM=Phi-2-2.7B, Res.=3842024.12 | 79.9 | |
| FastVModel=LLaVA-13B, TFLOPS=3.09, FLOPs Ratio=53%2025.03 | 79.9 | |
| Visual PromptLLM=Vicuna-7B, Res.=3362024.08 | 79.8 | |
| Bunny-3BLM=Phi-2-2.7B, Res.=3842024.12 | 79.8 | |
| BaselineModel=LLaVA-13B, TFLOPS=5.81, FLOPs Ratio=100%2025.03 | 79.8 | |
| S2-Wrapper*LLM=Vicuna-7B, Res.=10082024.08 | 79.7 | |
| HiMAPModel=InternVL-7B, TFLOPS=0.56, FLOPs Ratio=20%2025.03 | 79.6 | |
| Qwen-VLLLM=Qwen-7B, Res.=4482024.08 | 79.5 | |
| Imp-v1LM=Phi-2-2.7B, Res.=3842024.12 | 79.5 | |
| mPLUG-Owl2Vision Encoder=CLIP-L, Resolution=448, LLM=Llama-2-7B, Data=400M2025.01 | 79.4 | |
| VILA-UBackbone=LLaMA-2-7B, Resolution=384, Visual Input Type=Discrete2024.12 | 79.4 | |
| mPLUG-Owl2# Param=8B2024.08 | 79.4 | |
| mPLUG-Owl2Backbone=Llama2-7B, #Data=401M2024.06 | 79.4 | |
| mPLUG-Owl2LM=LLaMA-7B, Res.=4482024.12 | 79.4 | |
| BaselineModel=InternVL-7B, TFLOPS=2.71, FLOPs Ratio=100%2025.03 | 79.4 | |
| LLaMA-VIDLLM=Vicuna-7B, Res.=3362024.08 | 79.3 | |
| InternVL-7BBackbone=Vicuna-7B, #Data=>28.7B2024.06 | 79.3 | |
| Qwen-2.5-VL-7BBackbone=Qwen-2.5-VL-7B, Decoding Strategy=Standard2026.01 | 79.3 | |
| AlignGPTLLM=Vicuna-7B, Res.=3362024.08 | 79.1 | |
| LLaVA-1.5-7B + MMFuserLLM=Vicuna-7B, Training annotations observed=true2024.10 | 79.1 | |
| INAR-VLStrategy=INAR-VL2026.05 | 79.1 | |
| FastVModel=InternVL-7B, TFLOPS=1.39, FLOPs Ratio=52%2025.03 | 79 | |
| MAGICTarget Model=LLaVA-1.5-13B, Data Budget=20%2026.05 | 78.9 | |
| DoLABackbone=Qwen-2.5-VL-7B, Decoding Strategy=DoLA2026.01 | 78.85 | |
| LLaVA-1.5*LLM=Vicuna-7B, Res.=3362024.08 | 78.8 | |
| Qwen-VLLLM=Qwen-7B, Resolution=448, Pre-training samples=1.4B, Instruction tuning samples=50M2023.12 | 78.8 | |
| Qwen-VLLLM=Qwen-7B, Parameters=9.6B, Resolution=4482024.06 | 78.8 | |
| Qwen-VLLLM=Qwen-7B, Training annotations observed=true2024.10 | 78.8 | |
| Qwen-VLLLM=Qwen-7B2024.11 | 78.8 | |
| HiMAPModel=QwenVL-7B, TFLOPS=0.89, FLOPs Ratio=25%2025.03 | 78.8 | |
| Text-OnlyStrategy=Text-Only2026.05 | 78.8 | |
| HiMAPModel=LLaVA-7B, TFLOPS=0.73, FLOPs Ratio=24%2025.03 | 78.6 | |
| DoLABackbone=Qwen-2-VL-7B, Decoding Strategy=DoLA2026.01 | 78.57 | |
| LLaVA-1.5Vision Encoder=CLIP-L, Resolution=336, LLM=Vicuna-7B, Data=1.2M2025.01 | 78.5 |