Diagram Question Answering on AI2D
96.02AI2D AccuracyERNIE 5.0-Base
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| ERNIE 5.0-BaseModel type=pre-trained2026.02 | 96.02 | — | — | — | — | — | — | — | |
| Qwen3.5-27BMode=REASONING, Architecture=Dense, # Total Params=27B, # Activated Params=27B2026.04 | 92.9 | — | — | — | — | — | — | — | |
| Llama 3.2-11BLanguage Model Scale=11B, Training Data Scale=Large2025.01 | 91.1 | — | — | — | — | — | — | — | |
| InternVL3-78BSize=78B2025.12 | 89.7 | — | — | — | — | — | — | — | |
| GPT-5-highModel Category=Proprietary Models2026.06 | 89.7 | — | — | — | — | — | — | — | |
| GPT-5-highSize=-2026.06 | 89.7 | — | — | — | — | — | — | — | |
| MetaForgeTool Setting=w/ IID Tools2026.06 | 89.6 | — | — | — | — | — | — | — | |
| Keye-VL-1.5-8BSize=8B2025.12 | 89.5 | — | — | — | — | — | — | — | |
| Gemini-2.5-ProVariant=Pro 2.52025.12 | 89.5 | — | — | — | — | — | — | — | |
| GPT-52025.12 | 89.5 | — | — | — | — | — | — | — | |
| Qwen3-VL-32BSize=32B2025.12 | 89.5 | — | — | — | — | — | — | — | |
| Qwen3-VL-235B-A22BMode=Thinking, Architecture=MoE, # Total Params=236B, # Activated Params=23B2026.04 | 89.2 | — | — | — | — | — | — | — | |
| EXAONE 4.5 33BMode=REASONING, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 89 | — | — | — | — | — | — | — | |
| InternVL3-8B-MastersSize=8B2025.12 | 88.9 | — | — | — | — | — | — | — | |
| MastersBase Model=InternVL3-8B2025.12 | 88.9 | — | — | — | — | — | — | — | |
| Qwen3-VL-32BMode=Thinking, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 88.9 | — | — | — | — | — | — | — | |
| Qwen2.5-VL-72BSize=72B2025.12 | 88.7 | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B-MastersSize=7B2025.12 | 88.6 | — | — | — | — | — | — | — | |
| MastersBase Model=Qwen2.5-VL-7B2025.12 | 88.6 | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-MastersSize=8B2025.12 | 88.5 | — | — | — | — | — | — | — | |
| MastersBase Model=Qwen3-VL-8B2025.12 | 88.5 | — | — | — | — | — | — | — | |
| Qwen3-VL-235BToken Retention Ratio=100%2026.05 | 88.5 | — | — | — | — | — | — | — | |
| Gemini-2.5-ProModel Category=Proprietary Models2026.06 | 88.4 | — | — | — | — | — | — | — | |
| Gemini-2.5-ProSize=-2026.06 | 88.4 | — | — | — | — | — | — | — | |
| GPT-5-Mini2025.12 | 88.2 | — | — | — | — | — | — | — | |
| GPT-5 miniMode=REASONING: HIGH, Architecture=-, # Total Params=-, # Activated Params=-2026.04 | 88.2 | — | — | — | — | — | — | — | |
| MetaForgeTool Setting=w/ OOD Tools2026.06 | 88.15 | — | — | — | — | — | — | — | |
| GLM-4.5V2025.12 | 88.1 | — | — | — | — | — | — | — | |
| Qwen3-VL-4B-MastersSize=4B2025.12 | 88 | — | — | — | — | — | — | — | |
| GLM-4.1V-9BSize=9B2025.12 | 87.9 | — | — | — | — | — | — | — | |
| InternVL3.5-38BSize=38B2025.12 | 87.8 | — | — | — | — | — | — | — | |
| Qwen3-VL-32BToken Ratio=100%2026.05 | 87.63 | — | — | — | — | — | — | — | |
| VLsI-7BParameters=7B2024.12 | 87.3 | — | — | — | — | — | — | — | |
| InternVL3.5-8B-MastersSize=8B2025.12 | 87.2 | — | — | — | — | — | — | — | |
| MastersBase Model=InternVL3.5-8B2025.12 | 87.2 | — | — | — | — | — | — | — | |
| InternVL3.5-38BRatio=100%, Backbone=InternVL3.5-38B2026.05 | 87.05 | — | — | — | — | — | — | — | |
| RADIO1DTokens per frame/tile=256, TTFT (ms)=452.8, LLM Backbone=9B Nemotron2026.07 | 86.9 | — | — | — | — | — | — | — | |
| Keye-VL-8BSize=8B2025.12 | 86.7 | — | — | — | — | — | — | — | |
| Ovis2-8BSize=8B2025.12 | 86.6 | — | — | — | — | — | — | — | |
| F3AToken Retention Ratio=60%2026.05 | 86.43 | — | — | — | — | — | — | — | |
| SigLIP2-gTokens per frame/tile=256, TTFT (ms)=517.6, LLM Backbone=9B Nemotron2026.07 | 86.4 | — | — | — | — | — | — | — | |
| RADIO1DTokens per frame/tile=224, TTFT (ms)=410.2, LLM Backbone=9B Nemotron2026.07 | 86.4 | — | — | — | — | — | — | — | |
| SigLIP2-SO400mTokens per frame/tile=256, TTFT (ms)=440.5, LLM Backbone=9B Nemotron2026.07 | 86.2 | — | — | — | — | — | — | — | |
| MiniCPM-o2.6-8BSize=8B2025.12 | 86.1 | — | — | — | — | — | — | — | |
| C-RADIOv4-HTokens per frame/tile=256, TTFT (ms)=468.2, LLM Backbone=9B Nemotron2026.07 | 86 | — | — | — | — | — | — | — | |
| DivPruneToken Retention Ratio=60%2026.05 | 85.96 | — | — | — | — | — | — | — | |
| GPT-4.12025.12 | 85.9 | — | — | — | — | — | — | — | |
| Qwen2.5-VL-VIFParams=7B, Base model=Qwen-2.5-VL2026.05 | 85.78 | — | — | — | — | — | — | — | |
| VisionZipToken Retention Ratio=60%2026.05 | 85.72 | — | — | — | — | — | — | — | |
| Qwen3-VL-8BSize=8B2025.12 | 85.7 | — | — | — | — | — | — | — | |
| Ovis2-4BSize=4B2025.12 | 85.7 | — | — | — | — | — | — | — | |
| RADIO1DTokens per frame/tile=192, TTFT (ms)=380.4, LLM Backbone=9B Nemotron2026.07 | 85.7 | — | — | — | — | — | — | — | |
| LLaVA-OneVisionSize=72B2024.09 | 85.6 | — | — | — | — | — | — | — | |
| LLaVA-OneVision-72BSize=72B2025.12 | 85.6 | — | — | — | — | — | — | — | |
| TVI-CoTModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=8B2026.06 | 85.6 | — | — | — | — | — | — | — | |
| CDPrunerToken Retention Ratio=60%2026.05 | 85.41 | — | — | — | — | — | — | — | |
| Qwen2.5-VL-3B-MastersSize=3B2025.12 | 85.4 | — | — | — | — | — | — | — | |
| CoLTSize=8B, Backbone=Qwen3-VL-8B2026.06 | 85.4 | — | — | — | — | — | — | — | |
| PaliGemma2-3B + AuditDMResolution=448x448, Fine-tuning=per-task2025.12 | 85.3 | — | — | — | — | — | — | — | |
| F3ARatio=60%, Backbone=InternVL3.5-38B2026.05 | 85.22 | — | — | — | — | — | — | — | |
| InternVL3-8BSize=8B2025.12 | 85.2 | — | — | — | — | — | — | — | |
| NVLM-72BSize=72B2025.12 | 85.2 | — | — | — | — | — | — | — | |
| InternVL3-8BModel Category=Open-Source MLLMs, Model Scale=8B2026.06 | 85.2 | — | — | — | — | — | — | — | |
| InternVL3Size=8B2026.06 | 85.2 | — | — | — | — | — | — | — | |
| DivPruneRatio=60%, Backbone=InternVL3.5-38B2026.05 | 85.13 | — | — | — | — | — | — | — | |
| RADIO1DTokens per frame/tile=128, TTFT (ms)=373.2, LLM Backbone=9B Nemotron2026.07 | 85.1 | — | — | — | — | — | — | — | |
| RADIO1DTokens per frame/tile=64, TTFT (ms)=346.1, LLM Backbone=9B Nemotron2026.07 | 85 | — | — | — | — | — | — | — | |
| FastVToken Retention Ratio=60%2026.05 | 84.9 | — | — | — | — | — | — | — | |
| F3AToken Ratio=60%2026.05 | 84.84 | — | — | — | — | — | — | — | |
| CDPrunerToken Ratio=60%2026.05 | 84.72 | — | — | — | — | — | — | — | |
| PaliGemma2-28BResolution=448x448, Fine-tuning=per-task2025.12 | 84.6 | — | — | — | — | — | — | — | |
| GPT-4o2025.12 | 84.6 | — | — | — | — | — | — | — | |
| PaliGemma2-10BResolution=448x448, Fine-tuning=per-task2025.12 | 84.4 | — | — | — | — | — | — | — | |
| Claude-Opus-4.1Model Category=Proprietary Models2026.06 | 84.4 | — | — | — | — | — | — | — | |
| Claude-Opus-4.1Size=-2026.06 | 84.4 | — | — | — | — | — | — | — | |
| LLaVA-OneVision-1.5-8BSize=8B2025.12 | 84.2 | — | — | — | — | — | — | — | |
| BaselineBackbone=LLaVA-OneVision-1.5-8B-Instruct, Retain Tokens=100%, Compression Ratio=0%2026.02 | 84.2 | — | — | — | — | — | — | — | |
| VisionZipRatio=60%, Backbone=InternVL3.5-38B2026.05 | 84.2 | — | — | — | — | — | — | — | |
| LLaVA-OneVision-1.5-8BModel Category=Open-Source MLLMs, Model Scale=8B2026.06 | 84.2 | — | — | — | — | — | — | — | |
| LLaVA-OneVision-1.5Size=8B2026.06 | 84.2 | — | — | — | — | — | — | — | |
| Qwen3-VL-4BSize=4B2025.12 | 84.1 | — | — | — | — | — | — | — | |
| Qwen3-VL-4B-InstructParameters=4B2026.03 | 84.1 | — | — | — | — | — | — | — | |
| InternVL3.5-8BSize=8B2025.12 | 84 | — | — | — | — | — | — | — | |
| RADIO1DTokens per frame/tile=32, TTFT (ms)=335.1, LLM Backbone=9B Nemotron2026.07 | 84 | — | — | — | — | — | — | — | |
| FastVRatio=60%, Backbone=InternVL3.5-38B2026.05 | 83.92 | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BSize=7B2025.12 | 83.9 | — | — | — | — | — | — | — | |
| VanillaBackbone=Qwen3-VL-4B-FP8, Flops Ratio Reduction=0%, Precision=FP82026.04 | 83.9 | — | — | — | — | — | — | — | |
| InternVL2-8BParameters=8B2024.12 | 83.8 | — | — | — | — | — | — | — | |
| VanillaAverage Token Reduction=0.0%, Backbone=Qwen3-VL-8B2026.06 | 83.8 | — | — | — | — | — | — | — | |
| CDPrunerRatio=60%, Backbone=InternVL3.5-38B2026.05 | 83.74 | — | — | — | — | — | — | — | |
| Qwen2.5-VL-SFTParams=7B, Base model=Qwen-2.5-VL2026.05 | 83.67 | — | — | — | — | — | — | — | |
| LLaVA-OneVision-1.5-4BSize=4B2025.12 | 83.6 | — | — | — | — | — | — | — | |
| Qwen3-VL-8B (Baseline)Model Category=Baseline, Model Scale=8B2026.06 | 83.6 | — | — | — | — | — | — | — | |
| Qwen3-VL (Textual reasoning)Size=8B2026.06 | 83.6 | — | — | — | — | — | — | — | |
| MiMo-VL-8BSize=8B2025.12 | 83.5 | — | — | — | — | — | — | — | |
| VisionZipToken Retention Ratio=40%2026.05 | 83.44 | — | — | — | — | — | — | — | |
| F3AToken Retention Ratio=40%2026.05 | 83.42 | — | — | — | — | — | — | — | |
| Molmo-72BSize=72B2025.12 | 83.4 | — | — | — | — | — | — | — | |
| DivPruneToken Retention Ratio=40%2026.05 | 83.38 | — | — | — | — | — | — | — | |
| FastVToken Ratio=60%2026.05 | 83.32 | — | — | — | — | — | — | — |