Visual Question Answering on ChartQA (test)
90.8AccuracyClaude-3.5 Sonnet
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Claude-3.5 Sonnet2024.09 | 90.8 | — | |
| InternVL2-Llama-3-76B2024.09 | 88.4 | — | |
| Qwen2-VL-72B2024.09 | 88.3 | — | |
| Eagle2.5-8BAccess=Open weights only2026.03 | 87.5 | — | |
| MiniCPM-V-4.5-8BAccess=Open weights only2026.03 | 87.4 | — | |
| Molmo-72BNumber of crops=362024.09 | 87.3 | — | |
| Gemini 1.5 Pro2024.09 | 87.2 | — | |
| Kimi-VL-A3B-Instruct2026.03 | 87 | — | |
| MolmoPoint-8BAccess=MolmoPoint2026.03 | 86.8 | — | |
| InternVL3.5-8BAccess=Open weights only2026.03 | 86.7 | — | |
| Molmo2-4BAccess=Fully open2026.03 | 86.1 | — | |
| InternVL3.5-4BAccess=Open weights only2026.03 | 86 | — | |
| Molmo2-8BAccess=Fully open2026.03 | 86 | — | |
| GPT-4o-05132024.09 | 85.7 | — | |
| Llama-3.2V-90B-Instruct2024.09 | 85.5 | — | |
| PLM-8BAccess=Fully open2026.03 | 85.5 | — | |
| Gemini 1.5 Flash2024.09 | 85.4 | — | |
| Qwen3-VL-8BAccess=Open weights only2026.03 | 85.2 | — | |
| Qwen3-VL-4BAccess=Open weights only2026.03 | 85 | — | |
| Keye-VL-1.5-8BAccess=Open weights only2026.03 | 85 | — | |
| MolmoPoint-8B-O-7BAccess=Fully open2026.03 | 84.9 | — | |
| Qwen3-VL-32B-InstructContext Window=4K2026.03 | 84.3 | — | |
| PLM-3BAccess=Fully open2026.03 | 84.3 | — | |
| Molmo-7B-DNumber of crops=362024.09 | 84.1 | — | |
| Qwen3-VL-32B-InstructContext Window=32K2026.03 | 84 | — | |
| GPT-5Access=API call only2026.03 | 83.8 | — | |
| LLaVA OneVision-72Bdistilled=true2024.09 | 83.7 | — | |
| Llama-3.2V-11B-Instruct2024.09 | 83.4 | — | |
| InternVL2-8B2024.09 | 83.3 | — | |
| Phi-4-reasoning-vision-15B2026.03 | 83.3 | — | |
| Qwen3-VL-8B-InstructContext Window=32K2026.03 | 83.2 | — | |
| DyCo-RLBackbone=Qwen2.5-VL-3B, Zero-shot=true2026.06 | 83.2 | — | |
| Qwen3-VL-8B-InstructContext Window=4K2026.03 | 83.1 | — | |
| Qwen2-VL-7B2024.09 | 83 | — | |
| Open-Source SOTA (Yao et al., 2024)Model Type=Open-Source (<10B), Evaluation Protocol=VQA2025.01 | 82.4 | — | |
| IXC-2.5Model Type=Open-Source (<10B), Evaluation Protocol=VQA2025.01 | 82.2 | — | |
| GPT-5 miniAccess=API call only2026.03 | 82.1 | — | |
| Phi3.5-Vision-4B2024.09 | 81.8 | — | |
| Pixtral-12B2024.09 | 81.8 | — | |
| Claude-3 Haiku2024.09 | 81.7 | — | |
| Proprietary API SOTA (Megvii, 2024)Model Type=Proprietary API, Evaluation Protocol=VQA2025.01 | 81.2 | — | |
| GRPO BaselineBackbone=Qwen2.5-VL-3B, Zero-shot=true2026.06 | 81 | — | |
| GeminiOCR Usage=SOTA Overall2024.02 | 80.8 | — | |
| Claude-3 Opus2024.09 | 80.8 | — | |
| IXC-2.5-ChatModel Type=Open-Source (<10B), Evaluation Protocol=VQA2025.01 | 80.5 | — | |
| Molmo-7B-ONumber of crops=362024.09 | 80.4 | — | |
| Claude Sonnet 4.5Access=API call only2026.03 | 80.2 | — | |
| LLaVA OneVision-7Bdistilled=true2024.09 | 80 | — | |
| GPT-4V2024.09 | 78.1 | — | |
| MolmoE-1BNumber of crops=362024.09 | 78 | — | |
| Gemini 2.5 ProAccess=API call only2026.03 | 77.8 | — | |
| PaLI-3OCR Usage=Without OCR2024.02 | 77.3 | — | |
| Gemini 2.5 FlashAccess=API call only2026.03 | 76.8 | — | |
| ScreenAIOCR Usage=With OCR2024.02 | 76.7 | — | |
| ScreenAIOCR Usage=Without OCR2024.02 | 76.6 | — | |
| Phi-4-reasoning-vision-15BProtocol=force nothink2026.03 | 76.5 | — | |
| Cambrian-1-34Bdistilled=true2024.09 | 75.6 | — | |
| Cambrian-1-8Bdistilled=true2024.09 | 73.3 | — | |
| PaLI-XOCR pipeline input=true2023.05 | 72.3 | — | |
| PaLI-XOCR pipeline input=false2023.05 | 70.9 | — | |
| Method [8]OCR pipeline input=false2023.05 | 70.5 | — | |
| TILTOCR Usage=With OCR2024.02 | 70.4 | — | |
| GLM-4.1V-9BAccess=Open weights only2026.03 | 70 | — | |
| VisionTaPasData Table Setting=Gold Data Table Provided2022.03 | 61.84 | — | |
| xGen-MM-interleave-4Bdistilled=true2024.09 | 60 | — | |
| T5Data Table Setting=Gold Data Table Provided2022.03 | 59.8 | — | |
| VL-T5Data Table Setting=Gold Data Table Provided2022.03 | 59.12 | — | |
| Pix2Struct LargePretraining=Screenshot parsing, Pixel only=true2022.10 | 58.6 | — | |
| Pix2Struct BasePretraining=Screenshot parsing, Pixel only=true2022.10 | 56 | — | |
| TaPasData Table Setting=Gold Data Table Provided2022.03 | 51.8 | — | |
| Top Probability + Confidence ModulationModel=LaViDa-Instruct, Decoding Method=Top Probability with Suffix Anchored Confidence Modulation (Full Method), Prompting Strategy=zero-shot2026.05 | 45.92 | — | |
| VisionTaPasData Table Setting=Gold Data Table Not Provided2022.03 | 45.52 | — | |
| Method [46]OCR pipeline input=true2023.05 | 45.5 | — | |
| VTPtype=State of the art w/ pipelines2022.10 | 45.5 | — | |
| Top Margin + Confidence ModulationModel=LaViDa-Instruct, Decoding Method=Top Margin with Suffix Anchored Confidence Modulation (Full Method), Prompting Strategy=zero-shot2026.05 | 45.44 | — | |
| Top Margin + Suffix AnchorModel=LaViDa-Instruct, Decoding Method=Top Margin with Suffix Anchor, Prompting Strategy=zero-shot2026.05 | 45.08 | — | |
| Top Probability + Suffix AnchorModel=LaViDa-Instruct, Decoding Method=Top Probability with Suffix Anchor, Prompting Strategy=zero-shot2026.05 | 44.96 | — | |
| DonutPretraining=OCR, Pixel only=true2022.10 | 41.8 | — | |
| VL-T5Data Table Setting=Gold Data Table Not Provided2022.03 | 41.56 | — | |
| TaPasData Table Setting=Gold Data Table Not Provided2022.03 | 41.28 | — | |
| T5Data Table Setting=Gold Data Table Not Provided2022.03 | 41.04 | — | |
| gemma-3-12b-it2026.03 | 39 | — | |
| PlotQA*Data Table Setting=Gold Data Table Not Provided2022.03 | 38 | — | |
| MTVBackbone=LLaMA-3-8B, Method=Multimodal Task Vectors2024.06 | 34.9 | — | |
| PaliGemma-mix-3B2024.09 | 33.7 | — | |
| RandomModel=LaViDa-Instruct, Decoding Method=Random position selection, Prompting Strategy=zero-shot2026.05 | 27.12 | — | |
| 8-shot ICLBackbone=LLaMA-3-8B, Number of shots=8, Method=In-Context Learning2024.06 | 26.4 | — | |
| 4-shot ICLBackbone=LLaMA-3-8B, Number of shots=4, Method=In-Context Learning2024.06 | 25 | — | |
| Top ProbabilityModel=LaViDa-Instruct, Decoding Method=Top Probability (unmodified baseline), Prompting Strategy=zero-shot2026.05 | 24.12 | — | |
| Phi-4-mm-instruct2026.03 | 23.5 | — | |
| Top MarginModel=LaViDa-Instruct, Decoding Method=Top Margin (unmodified baseline), Prompting Strategy=zero-shot2026.05 | 23.24 | — | |
| 0-shot ICLBackbone=LLaMA-3-8B, Number of shots=0, Method=In-Context Learning2024.06 | 19.1 | — | |
| LLaVA-1.5-13B2024.09 | 18.2 | — | |
| LLaVA-1.5-7B2024.09 | 17.8 | — | |
| PREFILData Table Setting=Gold Data Table Not Provided2022.03 | 4.8 | — | |
| Donut2023.05 | — | 41.8 | |
| DUBLINresolution=fixed2023.05 | — | 35.6 | |
| DUBLINresolution=variable2023.05 | — | 35.2 | |
| Monkey2023.11 | — | 65.1 | |
| Pix2Structsize=large2023.05 | — | 58.6 |