Document Visual Question Answering on DocVQA (test)
96.5ANLSQwen-VL-Max-0809
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Qwen-VL-Max-0809Params (B)=722024.10 | 96.5 | — | — | — | — | — | |
| Qwen3-VL-8BParameter Scale=8B2026.05 | 96.1 | — | — | — | — | — | |
| Qwen3-VL-4BParameter Scale=4B2026.05 | 95.3 | — | — | — | — | — | |
| PerceptionLM-8BParameter Scale=8B2026.05 | 94.6 | — | — | — | — | — | |
| PerceptionLM-3BParameter Scale=3B2026.05 | 93.8 | — | — | — | — | — | |
| Qwen3-VL-2BParameter Scale=2B2026.05 | 93.3 | — | — | — | — | — | |
| DocCogitoSize(B)=82026.03 | 93.2 | — | — | — | — | — | |
| Molmo2-8BParameter Scale=8B2026.05 | 93.2 | — | — | — | — | — | |
| Qwen-VL-Maxopen-source=false2024.04 | 93.1 | — | — | — | — | — | |
| Gemini-1.5-ProModel Scale=High-End2024.09 | 93.1 | — | — | — | — | — | |
| DocCogitoSize(B)=42026.03 | 93.1 | — | — | — | — | — | |
| Zamba2-VL-7BParameter Scale=7B2026.05 | 92.9 | — | — | — | — | — | |
| GPT-4oModel Scale=High-End2024.09 | 92.8 | — | — | — | — | — | |
| InternVL3.5-4BParameter Scale=4B2026.05 | 92.4 | — | — | — | — | — | |
| InternVL3.5-8BParameter Scale=8B2026.05 | 92.3 | — | — | — | — | — | |
| SkillOptModel=GPT–5.5, Harness=Codex harness2026.05 | 92.2 | — | — | — | — | — | |
| MartenSize(B)=8.12026.03 | 92 | — | — | — | — | — | |
| InternVL2Size(B)=8.12026.03 | 91.6 | — | — | — | — | — | |
| Qwen3-VL-InstructSize(B)=82026.03 | 91.6 | — | — | — | — | — | |
| Qwen-VL-Plusopen-source=false2024.04 | 91.4 | — | — | — | — | — | |
| MM1.5-30BModel Scale=30B2024.09 | 91.4 | — | — | — | — | — | |
| SkillOptModel=Qwen3.6–35B-A3B, Harness=No harness / direct chat2026.05 | 91.4 | — | — | — | — | — | |
| SkillOptModel=GPT–5.5, Harness=No harness / direct chat2026.05 | 91.2 | — | — | — | — | — | |
| SkillOptModel=GPT–5.4, Harness=No harness / direct chat2026.05 | 91.2 | — | — | — | — | — | |
| Gemini Ultra 1.0open-source=false2024.04 | 90.9 | — | — | — | — | — | |
| InternVL 1.5#param=26B, open-source=false2024.04 | 90.9 | — | — | — | — | — | |
| GeminiOCR Usage=SOTA Overall2024.02 | 90.9 | — | — | — | — | — | |
| SkillOptModel=GPT–5.4-mini, Harness=No harness / direct chat2026.05 | 90.9 | — | — | — | — | — | |
| Zamba2-VL-2.7BParameter Scale=2.7B2026.05 | 90.9 | — | — | — | — | — | |
| SMoLA-PaLI-X_FT (Specialist)Base model=PaLI-X, Per-task LoRA tuning=true, LoRA rank=42023.12 | 90.8 | — | — | — | — | — | |
| PerceptionLM-1BParameter Scale=1B2026.05 | 90.7 | — | — | — | — | — | |
| SMoLA-PaLI-XFTOCR pipeline input=true, Model Category=Generalist, Backbone=PaLI-X2023.12 | 90.6 | — | — | — | — | — | |
| Trace2SkillModel=GPT–5.5, Harness=No harness / direct chat2026.05 | 90.6 | — | — | — | — | — | |
| LLM skillModel=GPT–5.4, Harness=No harness / direct chat2026.05 | 90.4 | — | — | — | — | — | |
| Trace2SkillModel=Qwen3.6–35B-A3B, Harness=No harness / direct chat2026.05 | 90.4 | — | — | — | — | — | |
| Qwen2-VLParams=2B2024.11 | 90.1 | — | — | — | — | — | |
| Human skillModel=GPT–5.5, Harness=No harness / direct chat2026.05 | 90.1 | — | — | — | — | — | |
| SkillOptModel=GPT–5.5, Harness=Claude Code harness2026.05 | 90.1 | — | — | — | — | — | |
| IXC2-4KHDModel Size=8B, Max Resolution=3840x16002024.04 | 90 | — | — | — | — | — | |
| ScreenAIOCR Usage=With OCR2024.02 | 89.9 | — | — | — | — | — | |
| LLM skillModel=GPT–5.5, Harness=Codex harness2026.05 | 89.8 | — | — | — | — | — | |
| TextHawk2Size(B)=7.42026.03 | 89.6 | — | — | — | — | — | |
| LLM skillModel=GPT–5.5, Harness=No harness / direct chat2026.05 | 89.6 | — | — | — | — | — | |
| SkillOptModel=GPT–5.2, Harness=No harness / direct chat2026.05 | 89.6 | — | — | — | — | — | |
| LLM skillModel=GPT–5.5, Harness=Claude Code harness2026.05 | 89.6 | — | — | — | — | — | |
| Claude-3 Sonnetopen-source=false2024.04 | 89.5 | — | — | — | — | — | |
| Claude 3 SonnetZero-shot=true2024.05 | 89.5 | — | — | — | — | — | |
| GEPAModel=GPT–5.2, Harness=No harness / direct chat2026.05 | 89.5 | — | — | — | — | — | |
| InternVL3.5-2BParameter Scale=2B2026.05 | 89.4 | — | — | — | — | — | |
| Claude-3 Opusopen-source=false2024.04 | 89.3 | — | — | — | — | — | |
| Hi-VT5OCR Usage=With OCR2024.02 | 89.3 | — | — | — | — | — | |
| Trace2SkillModel=GPT–5.4, Harness=No harness / direct chat2026.05 | 89.3 | — | — | — | — | — | |
| EvoSkillModel=GPT–5.5, Harness=Codex harness2026.05 | 89.3 | — | — | — | — | — | |
| InternVL2Params=4B2024.11 | 89.2 | — | — | — | — | — | |
| GEPAModel=GPT–5.5, Harness=No harness / direct chat2026.05 | 89.1 | — | — | — | — | — | |
| Qwen3-VL-InstructSize(B)=42026.03 | 89 | — | — | — | — | — | |
| Human skillModel=GPT–5.2, Harness=No harness / direct chat2026.05 | 89 | — | — | — | — | — | |
| SkillOptModel=Qwen3.5–4B, Harness=No harness / direct chat2026.05 | 89 | — | — | — | — | — | |
| Claude-3 Haikuopen-source=false2024.04 | 88.8 | — | — | — | — | — | |
| Claude 3 HaikuZero-shot=true2024.05 | 88.8 | — | — | — | — | — | |
| Human skillModel=GPT–5.5, Harness=Codex harness2026.05 | 88.8 | — | — | — | — | — | |
| PaLI-3 (Specialist)OCR pipeline input=true, Model Category=Specialist, Reference Citation=[8]2023.12 | 88.6 | — | — | — | — | — | |
| SOTA [8]2023.12 | 88.6 | — | — | — | — | — | |
| Human skillModel=GPT–5.4, Harness=No harness / direct chat2026.05 | 88.5 | — | — | — | — | — | |
| Trace2SkillModel=GPT–5.4-mini, Harness=No harness / direct chat2026.05 | 88.5 | — | — | — | — | — | |
| GPT-4Paradigm=Zero-shot, Parameters=Unknown, Text Modality=not clearly described, Vision Modality=true, Layout Modality=true, Fine-tuning Set=none2023.06 | 88.4 | — | — | — | — | — | |
| GPT-4Vopen-source=false2024.04 | 88.4 | — | — | — | — | — | |
| GPT-4VModel Scale=High-End2024.09 | 88.4 | — | — | — | — | — | |
| Human skillModel=Qwen3.6–35B-A3B, Harness=No harness / direct chat2026.05 | 88.4 | — | — | — | — | — | |
| LLM skillModel=Qwen3.6–35B-A3B, Harness=No harness / direct chat2026.05 | 88.4 | — | — | — | — | — | |
| GEPAModel=GPT–5.4, Harness=No harness / direct chat2026.05 | 88.3 | — | — | — | — | — | |
| Gemini Pro 1.0open-source=false2024.04 | 88.1 | — | — | — | — | — | |
| Gemini 1.0 ProZero-shot=true2024.05 | 88.1 | — | — | — | — | — | |
| MM1.5-7BModel Scale=7B2024.09 | 88.1 | — | — | — | — | — | |
| LLM skillModel=Qwen3.5–4B, Harness=No harness / direct chat2026.05 | 88 | — | — | — | — | — | |
| Trace2SkillModel=Qwen3.5–4B, Harness=No harness / direct chat2026.05 | 88 | — | — | — | — | — | |
| GEPAModel=Qwen3.6–35B-A3B, Harness=No harness / direct chat2026.05 | 88 | — | — | — | — | — | |
| Human skillModel=GPT–5.5, Harness=Claude Code harness2026.05 | 88 | — | — | — | — | — | |
| DocFormerv2largeModality=image + text + spatial features, Pre-train data=64M, #param=750M, Extra document VQA data=true2023.06 | 87.84 | — | — | — | — | — | |
| SmoLA PaLI-XOCR Usage=Without OCR2024.02 | 87.8 | — | — | — | — | — | |
| BlueLM-VParams=3B2024.11 | 87.8 | — | — | — | — | — | |
| DocFormerv2Size(B)=0.752026.03 | 87.8 | — | — | — | — | — | |
| Human skillModel=Qwen3.5–4B, Harness=No harness / direct chat2026.05 | 87.8 | — | — | — | — | — | |
| Molmo2-4BParameter Scale=4B2026.05 | 87.8 | — | — | — | — | — | |
| InternVL 1.2#param=40B, open-source=false2024.04 | 87.7 | — | — | — | — | — | |
| MM1.5-3BModel Scale=3B2024.09 | 87.7 | — | — | — | — | — | |
| LLM skillModel=GPT–5.2, Harness=No harness / direct chat2026.05 | 87.7 | — | — | — | — | — | |
| Trace2SkillModel=GPT–5.2, Harness=No harness / direct chat2026.05 | 87.7 | — | — | — | — | — | |
| PaLI-3 (Specialist)OCR pipeline input=false, Model Category=Specialist, Reference Citation=[8]2023.12 | 87.6 | — | — | — | — | — | |
| No skillModel=Qwen3.6–35B-A3B, Harness=No harness / direct chat2026.05 | 87.6 | — | — | — | — | — | |
| ScreenAIOCR Usage=Without OCR2024.02 | 87.5 | — | — | — | — | — | |
| SMoLA-PaLI-3FTOCR pipeline input=true, Model Category=Generalist, Backbone=PaLI-32023.12 | 87.4 | — | — | — | — | — | |
| Mini-MonkeySize(B)=22026.03 | 87.4 | — | — | — | — | — | |
| Zamba2-VL-1.2BParameter Scale=1.2B2026.05 | 87.4 | — | — | — | — | — | |
| DocFormerv2largeModality=image + text + spatial features, Pre-train data=64M, #param=750M, Extra document VQA data=false2023.06 | 87.2 | — | — | — | — | — | |
| TextGradModel=GPT–5.5, Harness=No harness / direct chat2026.05 | 87.2 | — | — | — | — | — | |
| TextGradModel=GPT–5.2, Harness=No harness / direct chat2026.05 | 87.2 | — | — | — | — | — | |
| No skillModel=GPT–5.5, Harness=Codex harness2026.05 | 87.2 | — | — | — | — | — | |
| EvoSkillModel=GPT–5.5, Harness=Claude Code harness2026.05 | 87.2 | — | — | — | — | — | |
| TILTSize variant=Large, Parameters=780M2021.02 | 87.05 | — | — | — | — | — |