Text-based Visual Question Answering on TextVQA (val)
86.5AccuracyPLM-8B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PLM-8BAccess=Fully open2026.03 | 86.5 | — | — | |
| MolmoPoint-8BAccess=MolmoPoint2026.03 | 86 | — | — | |
| Molmo2-8BAccess=Fully open2026.03 | 85.7 | — | — | |
| Qwen-VL-Max-0809Params (B)=722024.10 | 85.5 | — | — | |
| Molmo2-4BAccess=Fully open2026.03 | 85 | — | — | |
| Llama 3-V 405B2024.07 | 84.8 | — | — | |
| MolmoPoint-8B-O-7BAccess=Fully open2026.03 | 84.7 | — | — | |
| Qwen2-VLSize=7B, #token /image tile=Native resolution (unfixed)2024.12 | 84.3 | — | — | |
| Qwen2-VL# Parameters=8B2024.12 | 84.3 | — | — | |
| PLM-3BAccess=Fully open2026.03 | 84.3 | — | — | |
| Eagle2.5-8BAccess=Open weights only2026.03 | 83.7 | — | — | |
| Qwen2.5-VL-7B + SD-RPNThroughput=0.50×, Maximum visual tokens=5762026.04 | 83.5 | — | — | |
| Qwen2.5-VL-7B + Q-ZoomThroughput=0.81×, Maximum visual tokens=5762026.04 | 83.5 | — | — | |
| GPT-4o2024.12 | 83.4 | — | — | |
| jina-vlm2025.12 | 83.2 | — | — | |
| Llama 3-V 70B2024.07 | 83.1 | — | — | |
| Full PrecModel Backbone=Qwen2.5-VL-7B-Instruct, Quantization Method=Full Precision, Quantization Bit-width/Group-size=None, Zero-shot=true2025.08 | 83.05 | — | — | |
| Full PrecModel=Qwen2.5-VL-7B-Instruct, Bit-width=Full Precision2025.08 | 83.05 | — | — | |
| InternVL2# Parameters=40B2024.12 | 83 | — | — | |
| Qwen3-VL-8BAccess=Open weights only2026.03 | 82.8 | — | — | |
| DocCogitoSize(B)=82026.03 | 82.4 | — | — | |
| Gemini Ultra 1.0open-source=false2024.04 | 82.3 | — | — | |
| GPTAQModel Backbone=Qwen2.5-VL-7B-Instruct, Quantization Method=GPTAQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 82.25 | — | — | |
| MiniCPM-V-4.5-8BAccess=Open weights only2026.03 | 82.2 | — | — | |
| NVLM-D-1.0# Parameters=78B2024.12 | 82.1 | — | — | |
| SAIL-VL2Number of Parameters=2.7B2025.12 | 82.1 | — | — | |
| Full PrecModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=Full Precision, Quantization Bit-width/Group-size=None, Zero-shot=true2025.08 | 82.02 | — | — | |
| Full PrecModel=Qwen2-VL-7B-Instruct, Bit-width=Full Precision2025.08 | 82.02 | — | — | |
| SAIL-VL1.5Number of Parameters=2.5B2025.12 | 82 | — | — | |
| VLMQModel Backbone=Qwen2.5-VL-7B-Instruct, Quantization Method=VLMQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 81.82 | — | — | |
| GPTQModel Backbone=Qwen2.5-VL-7B-Instruct, Quantization Method=GPTQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 81.79 | — | — | |
| DocCogitoSize(B)=42026.03 | 81.6 | — | — | |
| Keye-VL-1.5-8BAccess=Open weights only2026.03 | 81.5 | — | — | |
| Qwen3-VL-4B + Q-ZoomThroughput=0.82×, Maximum visual tokens=5762026.04 | 81.4 | — | — | |
| VLMQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=VLMQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 81.21 | — | — | |
| ARIA2024.10 | 81.1 | — | — | |
| Qwen2.5-VL-7BThroughput=1.0×, Maximum visual tokens=5762026.04 | 81.1 | — | — | |
| GPTQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=GPTQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 81.02 | — | — | |
| Qwen3-VL-4BAccess=Open weights only2026.03 | 81 | — | — | |
| GPTAQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=GPTAQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 80.98 | — | — | |
| Qwen3-VL-InstructSize(B)=82026.03 | 80.7 | — | — | |
| InternVL 1.5#param=26B, open-source=false2024.04 | 80.6 | — | — | |
| LLaVA-OneVision# Parameters=72B2024.12 | 80.5 | — | — | |
| NVILA# Parameters=8B2024.12 | 80.1 | — | — | |
| PVC InternVL2 (Ours)Size=8B, #token /image tile=642024.12 | 80 | — | — | |
| NVILA# Parameters=15B2024.12 | 80 | — | — | |
| Qwen2.5-VL-3B + Q-ZoomThroughput=0.73×, Maximum visual tokens=5762026.04 | 80 | — | — | |
| AndesVLNumber of Parameters=2.4B2025.12 | 79.9 | — | — | |
| Qwen3-VLNumber of Parameters=2.1B2025.12 | 79.8 | — | — | |
| Qwen2-VL-2BParams (B)=2.12024.10 | 79.7 | — | — | |
| Qwen2-VLSize=2B, #token /image tile=Native resolution (unfixed)2024.12 | 79.7 | — | — | |
| Qwen2-VL-2BParameter Count=2B2025.12 | 79.7 | — | — | |
| Qwen2-VLNumber of Parameters=2.2B2025.12 | 79.7 | — | — | |
| Qwen2.5-VL-3B + SD-RPNThroughput=0.49×, Maximum visual tokens=5762026.04 | 79.7 | — | — | |
| GLM-4.1V-9BAccess=Open weights only2026.03 | 79.6 | — | — | |
| Qwen3-VL-2BParameter Count=2B2025.12 | 79.5 | — | — | |
| VLMQModel=Qwen2-VL-7B-Instruct, Bit-width=3-bit, Importance Factor Configuration=Default (ℓ1-norm)2025.08 | 79.41 | — | — | |
| Full PrecModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=Full Precision, Quantization Bit-width/Group-size=None, Zero-shot=true2025.08 | 79.38 | — | — | |
| Full PrecModel=Qwen2-VL-2B-Instruct, Bit-width=Full Precision2025.08 | 79.38 | — | — | |
| VLMQModel=Qwen2.5-VL-7B-Instruct, Bit-width=3-bit, Importance Factor Configuration=Default (ℓ1-norm)2025.08 | 79.28 | — | — | |
| GPTAQModel=Qwen2-VL-7B-Instruct, Bit-width=3-bit2025.08 | 79.23 | — | — | |
| Qwen3-VL-4BThroughput=1.0×, Maximum visual tokens=5762026.04 | 79.2 | — | — | |
| Qwen3-VL-InstructSize(B)=42026.03 | 79.1 | — | — | |
| GPT-5 miniAccess=API call only2026.03 | 79.1 | — | — | |
| AWQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=AWQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 78.91 | — | — | |
| IDEFICS-80BLLM=LLaMA-65B, Eff. Res.=224, PT=353M, SFT=1M2024.06 | 78.9 | — | — | |
| HyperVL ViTLNumber of Parameters=2.0B2025.12 | 78.8 | — | — | |
| Gemini 1.5 Pro2024.07 | 78.7 | — | — | |
| Gemini-1.5 Flash2024.10 | 78.7 | — | — | |
| Gemini-1.5 Pro2024.10 | 78.7 | — | — | |
| Gemini 1.5 Pro2024.12 | 78.7 | — | — | |
| GPT-5Access=API call only2026.03 | 78.7 | — | — | |
| MBQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=MBQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 78.67 | — | — | |
| GPTQModel=Qwen2.5-VL-7B-Instruct, Bit-width=3-bit2025.08 | 78.38 | — | — | |
| LLaVA-OneVision# Parameters=8B2024.12 | 78.3 | — | — | |
| GPTQModel=Qwen2-VL-7B-Instruct, Bit-width=3-bit2025.08 | 78.28 | — | — | |
| Llama 3-V 8B2024.07 | 78.2 | — | — | |
| IXC-2.5Size=7B, #token /image tile=4002024.12 | 78.2 | — | — | |
| InternVL3.5-8BAccess=Open weights only2026.03 | 78.2 | — | — | |
| Qwen3-VL-4B + SD-RPNThroughput=0.63×, Maximum visual tokens=5762026.04 | 78.2 | — | — | |
| Grok-1.5Vopen-source=false2024.04 | 78.1 | — | — | |
| NVILA-Lite# Parameters=8B2024.12 | 78.1 | — | — | |
| Ovis2Number of Parameters=2.5B2025.12 | 78.1 | — | — | |
| GPT-4Vopen-source=false2024.04 | 78 | — | — | |
| GPT-4V (2023.11.06)Proprietary=true2024.08 | 78 | — | — | |
| GPT-4V2024.07 | 78 | — | — | |
| GPT-4V2024.10 | 78 | — | — | |
| GPTAQModel=Qwen2.5-VL-7B-Instruct, Bit-width=3-bit2025.08 | 77.94 | — | — | |
| InternVL3.5-4BAccess=Open weights only2026.03 | 77.9 | — | — | |
| GPTAQModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=GPTAQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 77.81 | — | — | |
| VLMQModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=VLMQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 77.55 | — | — | |
| InternVL2Size=8B, #token /image tile=2562024.12 | 77.4 | — | — | |
| GPT-4o2024.12 | 77.4 | — | — | |
| InternVL2# Parameters=8B2024.12 | 77.4 | — | — | |
| InternVL2Size(B)=8.12026.03 | 77.4 | — | — | |
| NVILA-Lite# Parameters=15B2024.12 | 77.3 | — | — | |
| GPTQModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=GPTQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 77.23 | — | — | |
| Full PrecModel=Qwen2.5-VL-32B-Instruct, Bit-width=Full Precision2025.08 | 77.06 | — | — | |
| InternVL3-2BParameter Count=2B2025.12 | 77 | — | — | |
| InternVL3Number of Parameters=2.1B2025.12 | 77 | — | — |