Text-based Visual Question Answering on TextVQA (VQA^T)
78AccuracyGPT-4V
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4VLanguage Model Scale=Large, Training Data Scale=Large2025.01 | 78 | |
| Cambrian-1-34BLanguage Model Scale=34B, Training Data Scale=Large2025.01 | 76.7 | |
| CogVLM-17BPT tks.=256, Parm.=10B, Training set observed=false2024.07 | 70.4 | |
| MGM-HD+CALLLM=Vicuna-13B, Resolution setting=High resolution, Without OCR tokens evaluation protocol=true2024.05 | 68.8 | |
| LeoLanguage Model Scale=7B, Training Data Scale=Standard2025.01 | 68.8 | |
| VILA-1.5Language Model Scale=7B, Training Data Scale=Standard2025.01 | 68.5 | |
| SPHINX-MoELLM Training #P=8x7B, Input image resolution (Res.)=448, Pre-training samples (PT)=-, Instruction tuning samples (IT)=15.3M2024.08 | 68 | |
| MonkeyLanguage Model Scale=7B, Training Data Scale=Standard2025.01 | 67.6 | |
| MGM-HDLLM=Vicuna-13B, Resolution setting=High resolution, Without OCR tokens evaluation protocol=true2024.05 | 67.2 | |
| MGM-HD+CALLLM=Vicuna-7B, Resolution setting=High resolution, Without OCR tokens evaluation protocol=true2024.05 | 67.1 | |
| LLaVA-NeXT+CALLLM=Vicuna-13B, Resolution setting=High resolution, Without OCR tokens evaluation protocol=true2024.05 | 67.1 | |
| LLaVA-NeXTLLM=Vicuna-13B, Resolution setting=High resolution, Without OCR tokens evaluation protocol=true2024.05 | 67 | |
| VILA-13BPT tks.=576, Parm.=13B, Training set observed=false2024.07 | 66.6 | |
| Mini-Gemini-13BPT tks.=576, Parm.=13B, Training set observed=false2024.07 | 65.9 | |
| MGM-HDLLM=Vicuna-7B, Resolution setting=High resolution, Without OCR tokens evaluation protocol=true2024.05 | 65.5 | |
| Mini-Gemini-7BPT tks.=576, Parm.=7B, Training set observed=false2024.07 | 65.2 | |
| CoS-7BPT tks.=80, Parm.=532M, Training set observed=false2024.07 | 65.1 | |
| LLaVA-NeXT+CALLLM=Vicuna-7B, Resolution setting=High resolution, Without OCR tokens evaluation protocol=true2024.05 | 65 | |
| VILA-7BPT tks.=576, Parm.=7B, Training set observed=false2024.07 | 64.4 | |
| VILALanguage Model Scale=7B, Training Data Scale=Standard2025.01 | 64.4 | |
| LLaVA-NeXTLLM=Vicuna-7B, Resolution setting=High resolution, Without OCR tokens evaluation protocol=true2024.05 | 64.2 | |
| Qwen-VLPT tks.=256, Parm.=8B, Training set observed=false2024.07 | 63.8 | |
| Qwen-VL-7BLLM Training #P=7B, Input image resolution (Res.)=448, Pre-training samples (PT)=1.4B, Instruction tuning samples (IT)=50M2024.08 | 63.8 | |
| MGM+CALLLM=Vicuna-13B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 63.8 | |
| Qwen-VLLLM=Qwen-7B2025.12 | 63.8 | |
| VanillaBackbone=LLaVA-1.5-13B, Vision token count=1442026.03 | 63.8 | |
| MGM+CALLLM=Vicuna-7B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 63 | |
| MGMLLM=Vicuna-13B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 62.6 | |
| Qwen-VL-7B-ChatLLM Training #P=7B, Input image resolution (Res.)=448, Pre-training samples (PT)=1.4B, Instruction tuning samples (IT)=50M2024.08 | 61.5 | |
| LLaVA-1.5-13BPT tks.=576, Parm.=13B, Training set observed=false2024.07 | 61.3 | |
| LLaVA-1.5LLM Training #P=13B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=665K2024.08 | 61.3 | |
| MGMLLM=Vicuna-7B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 61.1 | |
| QMoPBackbone=LLaVA-1.5-13B, Vision token count=1442026.03 | 59.4 | |
| FasterVLMBackbone=LLaVA-1.5-13B, Vision token count=1442026.03 | 58.8 | |
| TokenPackerBackbone=LLaVA-1.5-13B, Vision token count=1442026.03 | 58.8 | |
| MoExtendLLM Training #P=3B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=665K2024.08 | 58.7 | |
| HyperLLaVALLM Training #P=7B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=665K2024.08 | 58.5 | |
| LDP-v2Backbone=LLaVA-1.5-13B, Vision token count=1442026.03 | 58.4 | |
| LLaVA-1.5-7BPT tks.=576, Parm.=7B, Training set observed=false2024.07 | 58.2 | |
| mPLUG-Owl2PT tks.=64, Parm.=7B, Training set observed=false2024.07 | 58.2 | |
| LLaVA-1.5LLM Training #P=7B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=665K2024.08 | 58.2 | |
| LLaVA-1.5LLM=Vicuna-7B2025.12 | 58.2 | |
| LLaVA.1.5Language Model Scale=7B, Training Data Scale=Standard2025.01 | 58.2 | |
| VanillaBackbone=LLaVA-1.5-7B, Vision token count=1442026.03 | 58.2 | |
| MoELoRABase Model=LLaVA-1.5-7B, Rank (r)=82024.08 | 57.1 | |
| TeamLoRABase Model=LLaVA-1.5-7B, Rank (r)=82024.08 | 57.1 | |
| FasterVLMBackbone=LLaVA-1.5-7B, Vision token count=1442026.03 | 57.1 | |
| QMoPBackbone=LLaVA-1.5-7B, Vision token count=1442026.03 | 57.1 | |
| MoE-LLaVA-2.7Bx4LLM Training #P=5B, Input image resolution (Res.)=384, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=1.6M2024.08 | 57 | |
| LLaVA-13bTrain #Tokens=18432, CR=100%, Time=21.1h2024.06 | 57 | |
| InternVLLanguage Model Scale=7B, Training Data Scale=Standard2025.01 | 57 | |
| TokenPackerBackbone=LLaVA-1.5-7B, Vision token count=1442026.03 | 57 | |
| Pixel-ShuffleBackbone=LLaVA-1.5-13B, Vision token count=1442026.03 | 57 | |
| LoRABase Model=LLaVA-1.5-7B, Rank (r)=322024.08 | 56.9 | |
| LLaVoltaTrain #Tokens=10863, CR=170%, Time=17.6h2024.06 | 56.7 | |
| LDP-v2Backbone=LLaVA-1.5-7B, Vision token count=1442026.03 | 56.4 | |
| FasterVLMBackbone=LLaVA-1.5-7B, Vision token count=642026.03 | 56 | |
| QMoPBackbone=LLaVA-1.5-7B, Vision token count=642026.03 | 56 | |
| TokenPackerBackbone=LLaVA-1.5-7B, Vision token count=642026.03 | 55.6 | |
| LDP-v2Backbone=LLaVA-1.5-7B, Vision token count=642026.03 | 55.1 | |
| Pixel-ShuffleBackbone=LLaVA-1.5-7B, Vision token count=1442026.03 | 54.2 | |
| C-AbstractorBackbone=LLaVA-1.5-7B, Vision token count=1442026.03 | 54.1 | |
| C-AbstractorBackbone=LLaVA-1.5-13B, Vision token count=1442026.03 | 53.9 | |
| MQT-LLaVABackbone=LLaVA-1.5-7B, Vision token count=1442026.03 | 52.6 | |
| Pixel-ShuffleBackbone=LLaVA-1.5-7B, Vision token count=642026.03 | 52.2 | |
| MGM+CALLLM=Gemma-2B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 51.8 | |
| MQT-LLaVABackbone=LLaVA-1.5-7B, Vision token count=642026.03 | 51.6 | |
| MoE-LLaVA-2.7Bx4LLM Training #P=5B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=1.6M2024.08 | 51.4 | |
| C-AbstractorBackbone=LLaVA-1.5-7B, Vision token count=642026.03 | 51.3 | |
| InstructBLIP-13BPT tks.=32, Parm.=188M, Training set observed=false2024.07 | 50.7 | |
| InstructBLIP-13BLLM Training #P=13B, Input image resolution (Res.)=224, Pre-training samples (PT)=129M, Instruction tuning samples (IT)=1.2M2024.08 | 50.7 | |
| InstructBLIPLLM=Vicuna-13B2025.12 | 50.7 | |
| InstructBLIP-7BPT tks.=32, Parm.=188M, Training set observed=false2024.07 | 50.1 | |
| InstructBLIP-7BLLM Training #P=7B, Input image resolution (Res.)=224, Pre-training samples (PT)=129M, Instruction tuning samples (IT)=1.2M2024.08 | 50.1 | |
| InstructBLIPLLM=Vicuna-7B2025.12 | 50.1 | |
| Instruct-BLIPLanguage Model Scale=7B, Training Data Scale=Standard2025.01 | 50.1 | |
| LLaVA-1.5+CALLLM=Vicuna-13B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 49.6 | |
| LLaVA-1.5LLM=Vicuna-13B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 49 | |
| MobileVLM 3BLLM=MobileLLaMA 2.7B, Tuning=Full-parameter2025.12 | 48.58 | |
| MGMLLM=Gemma-2B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 48.1 | |
| LLaVA-1.5LLM=Vicuna-7B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 47.6 | |
| LLaVA-1.5+CALLLM=Vicuna-7B, Resolution setting=Low resolution, Without OCR tokens evaluation protocol=true2024.05 | 47.5 | |
| HyDRA(Coarse-grained)LLM=MobileLLaMA 2.7B2025.12 | 47.14 | |
| HyDRA(Fine-grained)LLM=MobileLLaMA 2.7B2025.12 | 46.91 | |
| MobileVLM 3BLLM=MobileLLaMA 2.7B, Tuning=LoRA2025.12 | 46.58 | |
| BLIP-2LLM Training #P=13B, Input image resolution (Res.)=224, Pre-training samples (PT)=129M, Instruction tuning samples (IT)=-2024.08 | 42.5 | |
| BLIP-2LLM=Vicuna-13B2025.12 | 42.5 | |
| HyDRA(Coarse-grained)LLM=MobileLLaMA 1.4B2025.12 | 40.98 | |
| HyDRA(Fine-grained)LLM=MobileLLaMA 1.4B2025.12 | 40.75 | |
| MobileVLM 1.7BLLM=MobileLLaMA 1.4B, Tuning=Full-parameter2025.12 | 40.36 | |
| MobileVLM 1.7BLLM=MobileLLaMA 1.4B, Tuning=LoRA2025.12 | 40.28 | |
| IDEFICS-80BLLM Training #P=65B, Input image resolution (Res.)=224, Pre-training samples (PT)=353M, Instruction tuning samples (IT)=1M2024.08 | 30.9 | |
| IDEFICS-80BLLM=LLaMA-65B2025.12 | 30.9 | |
| IDEFICS-9BPT tks.=64, Parm.=9B, Training set observed=false2024.07 | 25.9 | |
| IDEFICS-9BLLM Training #P=7B, Input image resolution (Res.)=224, Pre-training samples (PT)=353M, Instruction tuning samples (IT)=1M2024.08 | 25.9 | |
| IDEFICS-9BLLM=LLaMA-7B2025.12 | 25.9 |