Visual Question Answering on TextVQA (val)
7,040VQA ScoreCogVLM-Chat
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| CogVLM-ChatLLM=Vicuna-7B, Trained during SFT stage=true2023.11 | 7,040 | — | — | — | |
| Emu2-ChatLLM=LLaMA-33B, Trained during SFT stage=true2023.11 | 6,660 | — | — | — | |
| Qwen-VL-ChatLLM=Qwen-7B, Trained during SFT stage=true2023.11 | 6,150 | — | — | — | |
| LLaVA-1.5LLM=Vicuna-13B, Trained during SFT stage=true2023.11 | 6,130 | — | — | — | |
| SPHINX-2kLLM=LLaMA2 13B, Trained during SFT stage=true2023.11 | 6,120 | — | — | — | |
| LLaVA-1.5LLM=Vicuna-7B, Trained during SFT stage=true2023.11 | 5,820 | — | — | — | |
| mPLUG-Owl2LLM=LLaMA2-7B, Trained during SFT stage=true2023.11 | 5,820 | — | — | — | |
| InstructBLIPLLM=Vicuna-7B, Trained during SFT stage=false2023.11 | 5,010 | — | — | — | |
| DreamLLMLLM=Vicuna-7B, Trained during SFT stage=false2023.11 | 3,490 | — | — | — | |
| IDEFICS-InstructLLM=LLaMA-65B, Trained during SFT stage=false2023.11 | 3,290 | — | — | — | |
| OpenFlamingoLLM=MPT-7B, Trained during SFT stage=false2023.11 | 2,830 | — | — | — | |
| Oracle2026.03 | 94.18 | — | — | — | |
| LCStraining_mode=learned, evaluation_protocol=5-fold cross-validation2026.03 | 86.48 | — | — | — | |
| QualRCCVrho=0.4, gamma=1, training_mode=training-free2026.03 | 86.07 | — | — | — | |
| RCCVrho=0.4, training_mode=training-free2026.03 | 85.97 | — | — | — | |
| Calibrated votepool_size=17 models, families=8 families2026.03 | 85.87 | — | — | — | |
| Majority votepool_size=17 models, families=8 families2026.03 | 85.67 | — | — | — | |
| HFVtraining_mode=training-free2026.03 | 85.27 | — | — | — | |
| HFV-sharptraining_mode=training-free2026.03 | 85.27 | — | — | — | |
| Human2021.05 | 85.01 | — | — | — | |
| FAAR-learntraining_mode=learned, evaluation_protocol=5-fold cross-validation2026.03 | 85 | — | — | — | |
| Qwen2.5-VL-7BPruning ratio=0%2026.03 | 84.8 | — | 100 | — | |
| Qwen2 VLInference mode=Zero-shot, Parameters=7B2024.10 | 84.3 | — | — | — | |
| Qwen2-VLSize=8B2026.01 | 84.3 | — | — | — | |
| Qwen2-VL-7BParameters=7B2026.06 | 84.3 | — | — | — | |
| Qwen3-OmniSize=30B-A3B, mode=instruct2026.04 | 84.1 | — | — | — | |
| MiniCPM-o 4.5Size=9B, mode=instruct2026.04 | 83.8 | — | — | — | |
| Qwen2.5-VL-7BSetup=Zero-shot, Base Model=Qwen2.5-VL-7B2026.03 | 82.9 | — | — | 100 | |
| Qwen3-VLSize=8B, mode=instruct2026.04 | 82.9 | — | — | — | |
| Single bestpool_size=17 models, families=8 families2026.03 | 82.88 | — | — | — | |
| Qwen3-VL-8BSetup=Zero-shot, Base Model=Qwen3-VL-8B2026.03 | 82.1 | — | — | 100 | |
| Proprietary API SOTA (Megvii, 2024)Model Type=Proprietary API, Evaluation Protocol=VQA2025.01 | 82 | — | — | — | |
| GPTAQModel=Qwen2-VL-7B-Instruct, Quantization Bit=INT42025.08 | 81.68 | — | — | — | |
| GPTQModel=Qwen2-VL-7B-Instruct, Quantization Bit=INT42025.08 | 81.55 | — | — | — | |
| VLMQModel=Qwen2-VL-7B-Instruct, Quantization Bit=INT42025.08 | 81.48 | — | — | — | |
| IXC-2.5-ChatModel Type=Open-Source (<10B), Evaluation Protocol=VQA2025.01 | 81.3 | — | — | — | |
| VisionZipSetup=Zero-shot, Base Model=Qwen2.5-VL-7B2026.03 | 81.3 | — | — | 50 | |
| ToMeSetup=Zero-shot, Base Model=Qwen2.5-VL-7B2026.03 | 81.2 | — | — | 50 | |
| Zamba2-VL-7BLanguage-backbone scale=7–8B2026.05 | 81 | — | — | — | |
| Qwen3-OmniInference mode=Thinking mode, Size=30B-A3B2026.04 | 80.8 | — | — | — | |
| PaLI-XOCR pipeline input=true2023.05 | 80.78 | — | — | — | |
| ResPrunePruning ratio=66.7%2026.03 | 80.3 | — | 98.4 | — | |
| NVILASize=8B2026.01 | 80.1 | — | — | — | |
| DivPrunePruning ratio=66.7%2026.03 | 80.1 | — | 96.7 | — | |
| Qwen2-VLParams=2B2024.11 | 79.9 | — | — | — | |
| MiniCPM-o 4.5Inference mode=Thinking mode, Size=9B2026.04 | 79.8 | — | — | — | |
| Qwen2.5VL-3BComplexity=O(n)2025.12 | 79.6 | — | — | — | |
| ToMeSetup=Zero-shot, Base Model=Qwen3-VL-8B2026.03 | 79.4 | — | — | 50 | |
| LinMU-NVSize=8B2026.01 | 79.3 | — | — | — | |
| VisionZipSetup=Zero-shot, Base Model=Qwen3-VL-8B2026.03 | 79.3 | — | — | 50 | |
| MM1.5-30BModel Scale=30B2024.09 | 79.2 | — | — | — | |
| Qwen-VL-PlusLLM=Private, Zero-shot=true2024.03 | 78.9 | — | — | — | |
| Gemini-1.5-ProModel Scale=High-End2024.09 | 78.7 | — | — | — | |
| Open-Source SOTA (Li et al., 2024a)Model Type=Open-Source (<10B), Evaluation Protocol=VQA2025.01 | 78.5 | — | — | — | |
| InfiniteVL-4BComplexity=O(1)2025.12 | 78.5 | — | — | — | |
| BlueLM-VParams=3B2024.11 | 78.4 | — | — | — | |
| LLaVA-OVSize=7B2026.01 | 78.3 | — | — | — | |
| IXC-2.5Model Type=Open-Source (<10B), Evaluation Protocol=VQA2025.01 | 78.2 | — | — | — | |
| CogVLMLLM=Vicuna-7B, #Sample=1500M, #Param=≥7B2026.01 | 78.2 | — | — | — | |
| InternVL3.5Size=8B, mode=instruct2026.04 | 78.2 | — | — | — | |
| PreciseDoc-Reasoner2026.06 | 78.2 | — | — | — | |
| Random DropSetup=Zero-shot, Base Model=Qwen2.5-VL-7B2026.03 | 78.1 | — | — | 50 | |
| GPT-4VLLM=Private, Zero-shot=true2024.03 | 78 | — | — | — | |
| ResPrunePruning ratio=77.8%2026.03 | 78 | — | 96.9 | — | |
| FastVPruning ratio=66.7%2026.03 | 77.9 | — | 92.3 | — | |
| GPT-5Inference mode=Thinking mode2026.04 | 77.8 | — | — | — | |
| Qwen3-VLInference mode=Thinking mode, Size=8B2026.04 | 77.8 | — | — | — | |
| CogVLM2025.03 | 77.57 | — | — | — | |
| CogVLM(base)Model Type=Original2025.03 | 77.57 | — | — | — | |
| GPT-4oSize=—2026.01 | 77.4 | — | — | — | |
| InternVL2Size=8B2026.01 | 77.4 | — | — | — | |
| Zamba2-VL-2.7BLanguage-backbone scale=2–4B2026.05 | 77.4 | — | — | — | |
| IXC2-4KHDModel Size=8B, Max Resolution=3840x16002024.04 | 77.2 | — | — | — | |
| MetaGPTunsupervised=true2025.03 | 77.18 | — | — | — | |
| Qwen3VL-8BParameters=8B2026.06 | 77 | — | — | — | |
| Qwen3.5-9BParameters=9B2026.06 | 77 | — | — | — | |
| AdaMMSunsupervised=true2025.03 | 76.89 | — | — | — | |
| MM1.5-3B-MOEModel Scale=3B, Architecture=MoE2024.09 | 76.8 | — | — | — | |
| InternVL2.5-4BComplexity=O(n)2025.12 | 76.8 | — | — | — | |
| Cambrian-34BModel Scale=34B2024.09 | 76.7 | — | — | — | |
| MiniCPM-Llama3-V 2.5Inference mode=Zero-shot, Parameters=8B2024.10 | 76.6 | — | — | — | |
| DivPrunePruning ratio=77.8%2026.03 | 76.6 | — | 95.1 | — | |
| Random DropSetup=Zero-shot, Base Model=Qwen3-VL-8B2026.03 | 76.6 | — | — | 50 | |
| MM1.5-3BModel Scale=3B2024.09 | 76.5 | — | — | — | |
| MM1.5-7BModel Scale=7B2024.09 | 76.5 | — | — | — | |
| CogAgentModel Size=17B, Max Resolution=1024x10242024.04 | 76.1 | — | — | — | |
| MM1.5-1B-MOEModel Scale=1B, Architecture=MoE2024.09 | 76.1 | — | — | — | |
| InternVL2Params=4B2024.11 | 74.7 | — | — | — | |
| Gemini ProLLM=Private, Zero-shot=true2024.03 | 74.6 | — | — | — | |
| Gemini 1.0 ProZero-shot=true2024.05 | 74.6 | — | — | — | |
| Baichuan-omniInference mode=Zero-shot, Parameters=7B2024.10 | 74.3 | — | — | — | |
| Gemini 2.5 FlashSize=-, mode=instruct2026.04 | 74.3 | — | — | — | |
| Mini-Gemini-HDLLM=Hermes-2-Yi-34B, Resolution=672, Zero-shot=true2024.03 | 74.1 | — | — | — | |
| Mini-Gemini-HDSize=34B, # tokens per image=2880, Zero-shot=true2024.05 | 74.1 | — | — | — | |
| MiniCPM-V 2.0-3BModel Scale=3B2024.09 | 74.1 | — | — | — | |
| Gemini 2.5 FlashInference mode=Thinking mode2026.04 | 73.8 | — | — | — | |
| Method [52]OCR pipeline input=true2023.05 | 73.67 | — | — | — | |
| MM1-ChatSize=30B, # tokens per image=720, Zero-shot=true2024.05 | 73.5 | — | — | — | |
| Gemini 1.5 ProZero-shot=true2024.05 | 73.5 | — | — | — | |
| MM1-30BModel Scale=30B2024.09 | 73.5 | — | — | — |