Document Visual Question Answering on DocVQA (val)
97.85AccuracyLLaVA-OV
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LLaVA-OVversion=1.5, parameters=8B2026.01 | 97.85 | — | |
| Qwen3-VLSize=8B, mode=instruct2026.04 | 96.1 | — | |
| Qwen3-VLparameters=8B2026.01 | 95.64 | — | |
| Qwen3-OmniSize=30B-A3B, mode=instruct2026.04 | 95.4 | — | |
| Innovator-VLvariant=8B-Instruct, parameters=8B2026.01 | 94.94 | — | |
| MiniCPM-Vversion=4.5, parameters=8B2026.01 | 94.94 | — | |
| Full PrecModel Backbone=Qwen2.5-VL-7B-Instruct, Quantization Method=Full Precision, Quantization Bit-width/Group-size=None, Zero-shot=true2025.08 | 94.72 | — | |
| Full PrecModel=Qwen2.5-VL-7B-Instruct, Bit-width=Full Precision2025.08 | 94.72 | — | |
| MiniCPM-o 4.5Size=9B, mode=instruct2026.04 | 94.7 | — | |
| Innovator-VLvariant=8B-Thinking, parameters=8B2026.01 | 94.56 | — | |
| Qwen2.5-VL-7B + Q-ZoomThroughput=0.81×, Maximum visual tokens=5762026.04 | 94.3 | — | |
| GPTQModel Backbone=Qwen2.5-VL-7B-Instruct, Quantization Method=GPTQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 93.91 | — | |
| Full PrecModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=Full Precision, Quantization Bit-width/Group-size=None, Zero-shot=true2025.08 | 93.89 | — | |
| Full PrecModel=Qwen2-VL-7B-Instruct, Bit-width=Full Precision2025.08 | 93.89 | — | |
| VLMQModel Backbone=Qwen2.5-VL-7B-Instruct, Quantization Method=VLMQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 93.83 | — | |
| GPTAQModel Backbone=Qwen2.5-VL-7B-Instruct, Quantization Method=GPTAQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 93.76 | — | |
| Qwen2-VLModel Size=8B2025.01 | 93.7 | — | |
| Qwen2.5-VL-7B + VisionThinkMaximum visual tokens=5762026.04 | 93.7 | — | |
| MBQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=MBQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 93.68 | — | |
| Qwen2.5-VL-7B + SD-RPNThroughput=0.50×, Maximum visual tokens=5762026.04 | 93.6 | — | |
| AWQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=AWQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 93.46 | — | |
| Qwen3-VL-4B + Q-ZoomThroughput=0.82×, Maximum visual tokens=5762026.04 | 93.4 | — | |
| VLMQModel=Qwen2-VL-7B-Instruct, Quantization Bit=INT42025.08 | 93.23 | — | |
| GPTQModel=Qwen2-VL-7B-Instruct, Quantization Bit=INT42025.08 | 93.16 | — | |
| GPTAQModel=Qwen2-VL-7B-Instruct, Quantization Bit=INT42025.08 | 93.04 | — | |
| Gemini 2.5 FlashSize=-, mode=instruct2026.04 | 93 | — | |
| Qwen3-VL-4B + SD-RPNThroughput=0.63×, Maximum visual tokens=5762026.04 | 92.8 | — | |
| Qwen2.5-VL-7B + AdaptVisionThroughput=0.06×, Maximum visual tokens=5762026.04 | 92.6 | — | |
| Intern-S1variant=mini, parameters=9B2026.01 | 92.59 | — | |
| Full PrecModel=Qwen2.5-VL-32B-Instruct, Bit-width=Full Precision2025.08 | 92.51 | — | |
| VLMQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=VLMQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 92.45 | — | |
| DeepSeekVL-2Model Size=8B2025.01 | 92.3 | — | |
| Qwen3-VL-2BParameter Count=2B, Evaluation Toolkit=VLMEvalKit2025.12 | 92.3 | — | |
| InternVL3.5Size=8B, mode=instruct2026.04 | 92.3 | — | |
| GPTAQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=GPTAQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 92.29 | — | |
| GPTQModel Backbone=Qwen2-VL-7B-Instruct, Quantization Method=GPTQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 92.21 | — | |
| SAIL-VLModel Size=8B2025.01 | 92.2 | — | |
| InternVL2.5-MPOModel Size=8B2025.01 | 92.1 | — | |
| Qwen2.5-VL-7BThroughput=1.0×, Maximum visual tokens=5762026.04 | 92 | — | |
| Qwen2.5-VL-3B + Q-ZoomThroughput=0.73×, Maximum visual tokens=5762026.04 | 91.8 | — | |
| LLaDA-oModel Architecture Style=Diff. Based2026.03 | 91.5 | — | |
| VLMQModel=Qwen2.5-VL-7B-Instruct, Bit-width=3-bit, Importance Factor Configuration=Default (ℓ1-norm)2025.08 | 91.36 | — | |
| Qwen3-VL-4BThroughput=1.0×, Maximum visual tokens=5762026.04 | 91.3 | — | |
| InternVL3.5parameters=8B2026.01 | 91.2 | — | |
| jina-vlm2025.12 | 90.6 | — | |
| GPTAQModel=Qwen2.5-VL-7B-Instruct, Bit-width=3-bit2025.08 | 90.41 | — | |
| InternVL-2-26BModel=InternVL-2-26B2026.01 | 90.4 | — | |
| IXC-2.5-7BModel=IXC-2.5-7B2026.01 | 90.3 | — | |
| InternVL-2-8BModel=InternVL-2-8B2026.01 | 90 | — | |
| Qwen2.5-VL-3B + SD-RPNThroughput=0.49×, Maximum visual tokens=5762026.04 | 89.5 | — | |
| Full PrecModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=Full Precision, Quantization Bit-width/Group-size=None, Zero-shot=true2025.08 | 89.35 | — | |
| Full PrecModel=Qwen2-VL-2B-Instruct, Bit-width=Full Precision2025.08 | 89.35 | — | |
| GPTQModel=Qwen2.5-VL-7B-Instruct, Bit-width=3-bit2025.08 | 89.33 | — | |
| SAIL-VLModel Size=2B2025.01 | 89.2 | — | |
| Qwen2-VL-2BParameter Count=2B, Evaluation Toolkit=VLMEvalKit2025.12 | 89.2 | — | |
| VLMQModel=Qwen2-VL-7B-Instruct, Bit-width=3-bit, Importance Factor Configuration=Default (ℓ1-norm)2025.08 | 88.9 | — | |
| GPTAQModel=Qwen2.5-VL-32B-Instruct, Bit-width=3-bit2025.08 | 88.72 | — | |
| Qwen2-VLModel Size=2B2025.01 | 88.7 | — | |
| DeepSeekVL-2Model Size=2B2025.01 | 88.6 | — | |
| InternVL3.5-2BParameter Count=2B, Evaluation Toolkit=VLMEvalKit2025.12 | 88.5 | — | |
| GPTQModel=Qwen2-VL-7B-Instruct, Bit-width=3-bit2025.08 | 88.35 | — | |
| GPTQModel=Qwen2.5-VL-32B-Instruct, Bit-width=3-bit2025.08 | 88.26 | — | |
| GPTAQModel=Qwen2-VL-7B-Instruct, Bit-width=3-bit2025.08 | 88.09 | — | |
| InternVL2.5-MPOModel Size=2B2025.01 | 87.8 | — | |
| VLMQModel=Qwen2.5-VL-32B-Instruct, Bit-width=3-bit, Importance Factor Configuration=Default (ℓ1-norm)2025.08 | 87.47 | — | |
| InternVL3-2BParameter Count=2B, Evaluation Toolkit=VLMEvalKit2025.12 | 87.4 | — | |
| Qwen2.5-VL-3BThroughput=1.0×, Maximum visual tokens=5762026.04 | 87.1 | — | |
| Full PrecModel Backbone=LLaVA-OneVision-7B, Quantization Method=Full Precision, Quantization Bit-width/Group-size=None, Zero-shot=true2025.08 | 87.09 | — | |
| Full PrecModel=LLaVA-OneVision-7B, Bit-width=Full Precision2025.08 | 87.09 | — | |
| MiMo-VLtraining=7B-RL, parameters=7B2026.01 | 87.05 | — | |
| MiMo-VLtraining=7B-SFT, parameters=7B2026.01 | 86.51 | — | |
| GPTQModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=GPTQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 85.29 | — | |
| VLMQModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=VLMQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 85.06 | — | |
| GPTAQModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=GPTAQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 84.92 | — | |
| GPTQModel Backbone=LLaVA-OneVision-7B, Quantization Method=GPTQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 84.53 | — | |
| GPTAQModel Backbone=LLaVA-OneVision-7B, Quantization Method=GPTAQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 84.06 | — | |
| AWQModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=AWQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 83.96 | — | |
| LLaDA-VModel Architecture Style=Diff. Based2026.03 | 83.9 | — | |
| VLMQModel Backbone=LLaVA-OneVision-7B, Quantization Method=VLMQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 83.89 | — | |
| LLaVA-OV-7B w/ CLIModel=LLaVA-OV-7B w/ CLI2026.01 | 82.8 | — | |
| LLaVA-OV-7BModel=LLaVA-OV-7B2026.01 | 82.5 | — | |
| MBQModel Backbone=Qwen2-VL-2B-Instruct, Quantization Method=MBQ, Quantization Bit-width/Group-size=INT3g128, Zero-shot=true2025.08 | 82.07 | — | |
| Qwen3-VL-8BFine-tuning setting=LoRA2026.07 | 79.92 | — | |
| LLaVA-OV-7B w/ SLIModel=LLaVA-OV-7B w/ SLI2026.01 | 79.9 | — | |
| GPTQBase Model=Qwen2.5-VL-7B-Instruct, Bit-width=2-bit, Group Size=1282025.08 | 78.79 | — | |
| VLMQModel=LLaVA-OneVision-7B, Bit-width=3-bit, Importance Factor Configuration=Default (ℓ1-norm)2025.08 | 78.32 | — | |
| VLMQModel=LLaVA-OneVision-7B, Bit-width=3-bit, Importance Factor Configuration=token-level ℓ2-norm2025.08 | 78.21 | — | |
| GPTAQModel=LLaVA-OneVision-7B, Bit-width=3-bit2025.08 | 78.07 | — | |
| GPTQModel=LLaVA-OneVision-7B, Bit-width=3-bit2025.08 | 77.82 | — | |
| Qwen2VL-7B-InstructLatency=0.35692025.12 | 77.38 | 93.9 | |
| CPOModel=Qwen3-VL-4B2026.07 | 77.08 | — | |
| VLMQModel=Qwen2-VL-2B-Instruct, Bit-width=3-bit, Importance Factor Configuration=Default (ℓ1-norm)2025.08 | 76.59 | — | |
| Qwen2VL-7B-Instruct + LUVCLatency=0.2335, Speedup=1.53x2025.12 | 76.52 | 91.33 | |
| Emu3Model Architecture Style=AR Based2026.03 | 76.3 | — | |
| VLMQBase Model=Qwen2-VL-7B-Instruct, Bit-width=2-bit, Group Size=128, Precursor Algorithm=GPTQ2025.08 | 75.76 | — | |
| LoRAModel=Qwen3-VL-4B2026.07 | 75.19 | — | |
| GPTQBase Model=Qwen2-VL-7B-Instruct, Bit-width=2-bit, Group Size=1282025.08 | 74.9 | — | |
| GPTAQBase Model=Qwen2.5-VL-7B-Instruct, Bit-width=2-bit, Group Size=1282025.08 | 74.82 | — | |
| GPTQModel=Qwen2-VL-2B-Instruct, Bit-width=3-bit2025.08 | 74.55 | — | |
| SigLIP SO400M + TilingVision Encoder=SigLIP SO400M [52] + Tiling, Resolution=Up to 13 × 384², Compression=2 × 2 Unshuffle, Tokens/im=~ 19282024.12 | 74.4 | — |