Document Visual Question Answering on InfoVQA
0.902AccuracyRLR³
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RLR³Training Source=OpenMMR2026.05 | 0.902 | — | |
| RLVRTraining Source=DeepVision2026.05 | 0.902 | — | |
| RLR³Training Source=DeepVision2026.05 | 0.895 | — | |
| RLR³Training Source=ViRL2026.05 | 0.893 | — | |
| RLVRTraining Source=ViRL2026.05 | 0.891 | — | |
| RLVRTraining Source=OpenMMR2026.05 | 0.873 | — | |
| Base instruct2026.05 | 0.867 | — | |
| AdaTooler-V-7BParameters=7B2025.12 | 0.86 | — | |
| Official thinkingSource=Qwen3-VL2026.05 | 0.856 | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Random2026.04 | 0.854 | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Complete2026.04 | 0.853 | — | |
| CompleteModel=Qwen3-VL 8B-Instruct2026.04 | 0.85 | — | |
| RandomModel=Qwen3-VL 8B-Instruct2026.04 | 0.849 | — | |
| SkillGraphModel=Qwen2.5-VL 7B-Instruct, Baseline=Random2026.04 | 0.846 | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Layered2026.04 | 0.845 | — | |
| SkillGraphModel=Qwen2.5-VL 7B-Instruct, Baseline=Complete2026.04 | 0.844 | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Linear2026.04 | 0.842 | — | |
| SkillGraphModel=Qwen2.5-VL 7B-Instruct, Baseline=Layered2026.04 | 0.842 | — | |
| CompleteModel=Qwen2.5-VL 7B-Instruct2026.04 | 0.842 | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Centralized2026.04 | 0.841 | — | |
| SkillGraphModel=Qwen2.5-VL 7B-Instruct, Baseline=Linear2026.04 | 0.841 | — | |
| SkillGraphModel=Qwen2.5-VL 7B-Instruct, Baseline=Centralized2026.04 | 0.841 | — | |
| RandomModel=Qwen2.5-VL 7B-Instruct2026.04 | 0.841 | — | |
| Pixel ReasonerCategory=Open-Source o3-like2025.12 | 0.84 | — | |
| DirectAnswerModel=Qwen3-VL 8B-Instruct2026.04 | 0.832 | — | |
| LayeredModel=Qwen3-VL 8B-Instruct2026.04 | 0.829 | — | |
| DirectAnswerModel=Qwen2.5-VL 7B-Instruct2026.04 | 0.829 | — | |
| CentralizedModel=Qwen3-VL 8B-Instruct2026.04 | 0.828 | — | |
| LayeredModel=Qwen2.5-VL 7B-Instruct2026.04 | 0.828 | — | |
| LinearModel=Qwen3-VL 8B-Instruct2026.04 | 0.827 | — | |
| Qwen2.5-VL-7B-InstructParameters=7B2025.12 | 0.826 | — | |
| CentralizedModel=Qwen2.5-VL 7B-Instruct2026.04 | 0.826 | — | |
| LinearModel=Qwen2.5-VL 7B-Instruct2026.04 | 0.825 | — | |
| Official instructSource=Qwen3-VL2026.05 | 0.818 | — | |
| Gemini 1.5 ProCategory=Proprietary2025.12 | 0.81 | — | |
| GPT-4oCategory=Proprietary2025.12 | 0.807 | — | |
| Qwen3-VL-4B-InstructParameters=4B2026.03 | 0.803 | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Complete2026.04 | 0.802 | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Random2026.04 | 0.791 | — | |
| DeFactoBackbone=Qwen2.5-VL-7B2025.09 | 0.791 | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Layered2026.04 | 0.789 | — | |
| CompleteModel=InternVL3-8B2026.04 | 0.787 | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Linear2026.04 | 0.786 | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Centralized2026.04 | 0.786 | — | |
| RandomModel=InternVL3-8B2026.04 | 0.785 | — | |
| GPT-5 miniVersion=high2026.05 | 0.776 | — | |
| DirectAnswerModel=InternVL3-8B2026.04 | 0.771 | — | |
| InternVL3-8BParameters=8B2025.12 | 0.768 | — | |
| LayeredModel=InternVL3-8B2026.04 | 0.768 | — | |
| CentralizedModel=InternVL3-8B2026.04 | 0.768 | — | |
| LinearModel=InternVL3-8B2026.04 | 0.766 | — | |
| Visual-SR1Backbone=Qwen2.5-VL-7B2025.09 | 0.765 | — | |
| dots.mocrParameters=3B2026.03 | 0.7376 | — | |
| SkillGraphModel=LLaVA-OV Qwen2-7B, Baseline=Random2026.04 | 0.736 | — | |
| SkillGraphModel=LLaVA-OV Qwen2-7B, Baseline=Complete2026.04 | 0.731 | — | |
| RandomModel=LLaVA-OV Qwen2-7B2026.04 | 0.73 | — | |
| SkillGraphModel=LLaVA-OV Qwen2-7B, Baseline=Linear2026.04 | 0.728 | — | |
| GPT-5 miniVersion=minimal2026.05 | 0.728 | — | |
| SkillGraphModel=LLaVA-OV Qwen2-7B, Baseline=Layered2026.04 | 0.727 | — | |
| CompleteModel=LLaVA-OV Qwen2-7B2026.04 | 0.727 | — | |
| SkillGraphModel=LLaVA-OV Qwen2-7B, Baseline=Centralized2026.04 | 0.726 | — | |
| Qwen3-VL-2B-InstructParameters=2B2026.03 | 0.724 | — | |
| LayeredModel=LLaVA-OV Qwen2-7B2026.04 | 0.715 | — | |
| Qwen2.5-VLBackbone=Qwen2.5-VL-7B2025.09 | 0.715 | — | |
| DirectAnswerModel=LLaVA-OV Qwen2-7B2026.04 | 0.713 | — | |
| CentralizedModel=LLaVA-OV Qwen2-7B2026.04 | 0.713 | — | |
| LinearModel=LLaVA-OV Qwen2-7B2026.04 | 0.711 | — | |
| LLaVA-OneVision-7BParameters=7B2025.12 | 0.688 | — | |
| ViCropBackbone=LLaVA-1.5 (Vicuna-7B)2025.09 | 0.524 | — | |
| GRITBackbone=Qwen2.5-VL-3B2025.09 | 0.437 | — | |
| DeepEyesBackbone=Qwen2.5-VL-7B2025.09 | 0.399 | — | |
| TextHarmony-ChatModel Generation Capability=Text and image, Slide-LoRA=true2024.07 | 0.289 | — | |
| InternLM-XComposer2Model Generation Capability=Text only2024.07 | 0.286 | — | |
| TextHarmonyModel Generation Capability=Text and image, Slide-LoRA=true2024.07 | 0.285 | — | |
| TextHarmony*Model Generation Capability=Text and image, Slide-LoRA=false2024.07 | 0.261 | — | |
| MonkeyModel Generation Capability=Text only2024.07 | 0.258 | — | |
| InternVLModel Generation Capability=Text only2024.07 | 0.236 | — | |
| SEED-LLaMA-14BModel Generation Capability=Text and image2024.07 | 0.235 | — | |
| mPLUG-Owl2Model Generation Capability=Text only2024.07 | 0.189 | — | |
| MM-InterleavedModel Generation Capability=Text and image2024.07 | 0.17 | — | |
| LLaVARModel Generation Capability=Text only2024.07 | 0.165 | — | |
| DocPediaModel Generation Capability=Text only2024.07 | 0.152 | — | |
| UniDocModel Generation Capability=Text only2024.07 | 0.147 | — | |
| LLaVA1.5-7BModel Generation Capability=Text only2024.07 | 0.147 | — | |
| MiniGPT5Model Generation Capability=Text and image2024.07 | 0.021 | — | |
| Cambrian-1LLM Backbone=LLaMA3-8B, Open Data=true, Super-Image (SI)=false2025.12 | — | 41.6 | |
| Claude 3.5 SonnetOpen Data=false, Super-Image (SI)=false2025.12 | — | 49.7 | |
| CogAgentSize=17B, Visual tokens=66562024.09 | — | 44.5 | |
| DeepSeek-VL2LLM Backbone=DeepSeek-MoE 27B, Open Data=false, Super-Image (SI)=false2025.12 | — | 61.3 | |
| DocOwlSize=7B, Visual tokens=8412024.09 | — | 38.2 | |
| DocOwl 1.5Size=8B, Visual tokens=16982024.09 | — | 50.7 | |
| DocOwl2Size=8B, Visual tokens=3242024.09 | — | 46.4 | |
| DocPediaSize=7B, Visual tokens=16002024.09 | — | 15.2 | |
| DonutSize=<1B, Visual tokens=4800, Fine-tuned=true2024.09 | — | 11.6 | |
| Dream-VLLLM Backbone=Dream 7B, Open Data=true, Super-Image (SI)=true2025.12 | — | 81.4 | |
| Dream-VLLLM Backbone=Dream 7B, Open Data=true, Super-Image (SI)=false2025.12 | — | 81 | |
| Gemini-1.5 ProOpen Data=false, Super-Image (SI)=false2025.12 | — | 81 | |
| GPT-4oOpen Data=false, Super-Image (SI)=false2025.12 | — | 79.2 | |
| InternVL 2Size=8B, Visual tokens=31332024.09 | — | 74.8 | |
| InternVL3LLM Backbone=Qwen2.5-7B, Open Data=false, Super-Image (SI)=false2025.12 | — | 76.8 |