Document Understanding on DocVQA
96.7ANLSOpenVLThinkerV2
Evaluation Results
| Method | Links | |
|---|---|---|
| OpenVLThinkerV22026.04 | 96.7 | |
| Qwen3-VL GRPOAlignment=GRPO2026.04 | 95.9 | |
| Qwen3-VL GDPOAlignment=GDPO2026.04 | 95.6 | |
| RegionDoc-R12026.04 | 95.3 | |
| Qwen3-VL-InstructModel Type=Instruct2026.04 | 95.3 | |
| VisionZero2026.04 | 95.2 | |
| OneThinker-8BParameters=8B2026.04 | 95 | |
| VisionThink2026.04 | 94.4 | |
| Gemini 2.5 Pro2026.04 | 92.6 | |
| VideoLLaMA3Architecture Category=Token Insertion – Proprietary2025.12 | 91.9 | |
| AVLM-2BModel Scale=2B2026.06 | 91.51 | |
| GPT-52026.04 | 91.5 | |
| Qwen2-VL# train tokens=1.4T, Architecture Category=Token Insertion – Proprietary2025.12 | 90.1 | |
| InsertionHe-2B# train tokens=0.1T, Architecture Category=Token Insertion – Public data, LLM Size=2B2025.12 | 89.1 | |
| InternVL2.5# train tokens=0.5T, Architecture Category=Token Insertion – Proprietary2025.12 | 88.7 | |
| CASA→# train tokens=0.3T, Architecture=CASAHe-2B, CASA design variant=→, LLM Size=2B2025.12 | 83.7 | |
| CASA⊕# train tokens=0.3T, Architecture=CASAHe-2B, CASA design variant=⊕, LLM Size=2B2025.12 | 82.8 | |
| CASA∨# train tokens=0.3T, Architecture=CASAHe-2B, CASA design variant=∨, LLM Size=2B2025.12 | 81.3 | |
| SmolVLMArchitecture Category=Token Insertion – Public data, LLM Size=2B2025.12 | 80 | |
| DocThinker-7BParameters=7B2026.04 | 78.8 | |
| Gemma-4-E2B-itModel Scale=4B2026.06 | 73.49 | |
| mPLUG-Owl3 8B# train tokens=0.1T, Architecture Category=Cross-attention-based – Public data, LLM Size=8B2025.12 | 55.9 | |
| mPLUG-Owl3 2B# train tokens=0.1T, Architecture Category=Cross-attention-based – Public data, LLM Size=2B2025.12 | 48.2 | |
| Qwen2.5-Omni-3BModel Scale=3B2026.06 | 38.27 |