Diagram Understanding on AI2D
94.2AccuracyGPT-4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4o2024.08 | 94.2 | |
| InternVL2.5-8B-GenRecalTeacher VLM=InternVL2.5-78B2025.06 | 93 | |
| LaVerVisual Encoder Model=Qwen-ViT2025.12 | 92.71 | |
| Qwen3.5Parameters=35BA3B, Internal Reasoning (Think mode)=false2026.05 | 92.6 | |
| SenseNova-U1Parameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 92.23 | |
| InternVL2.5-8B-GenRecalTeacher VLM=Qwen2-VL-72B2025.06 | 92.2 | |
| InternVL2.5-8B-GenRecalTeacher VLM=InternVL2-76B2025.06 | 92.1 | |
| SenseNova-U1Parameters=8B, Internal Reasoning (Think mode)=true2026.05 | 91.74 | |
| OmniInference scheme=no-thinking2026.04 | 91.5 | |
| Gemini 2.5 ProAccess=Close-source2026.02 | 90.9 | |
| Qwen3.5Parameters=9B, Internal Reasoning (Think mode)=false2026.05 | 90.2 | |
| Bagel-7B-MoTModel Category=Open-source General Models2025.11 | 89.5 | |
| GPT-5Inference mode=Thinking mode2026.04 | 89.5 | |
| InternVL2.5-78B2025.06 | 89.1 | |
| LaVerVisual Encoder Model=SigLIP22025.12 | 89.09 | |
| VLSI-2BParameters=2B2024.12 | 89 | |
| VLSI-2BSize=2B2024.12 | 89 | |
| SenseNova-SIBase Architecture=Bagel-7B-MoT, Model Category=Ours2025.11 | 88.8 | |
| Claude-3-SonnetTable Header=Claude-S2026.02 | 88.7 | |
| Gemini 2.5 FlashInference mode=Thinking mode2026.04 | 88.7 | |
| InternVL2.5-8B-GenRecalTeacher VLM=NVLM-72B2025.06 | 88.5 | |
| MiniCPM-o 4.5Inference mode=Thinking mode, Size=9B2026.04 | 88.5 | |
| Qwen2-VL-72BSize=72B2024.12 | 88.3 | |
| Grok-1.5V2026.02 | 88.3 | |
| Claude-3-OpusTable Header=Claude-O2026.02 | 88.1 | |
| Qwen2-VL-72B2025.06 | 88.1 | |
| BaselineVisual Encoder Model=Qwen-ViT2025.12 | 88 | |
| InternVL3.5-38BLLM=Qwen3-32B2025.12 | 87.8 | |
| Qwen2.5-VL-32BTraining=DF-GSPO2026.03 | 87.8 | |
| InternVL2-76BSize=76B2024.12 | 87.6 | |
| InternVL2-76B2025.06 | 87.6 | |
| LaVerVisual Encoder Model=AIMv22025.12 | 87.46 | |
| VLSI-7BSize=7B2024.12 | 87.3 | |
| Qwen3VLParameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 86.9 | |
| InternVL3.5-30B-A3BInference scheme=no-thinking2026.04 | 86.8 | |
| BaselineVisual Encoder Model=SigLIP22025.12 | 86.51 | |
| LLaVA-OV-72BSize=72B2024.12 | 86.2 | |
| R1-ShareVL-32B2026.03 | 86.2 | |
| Penguin-VLModel Size=8B2026.03 | 86.1 | |
| Qwen3-OmniInference mode=Thinking mode, Size=30B-A3B2026.04 | 86.1 | |
| Gemma4Parameters=26BA4B, Internal Reasoning (Think mode)=false2026.05 | 86.04 | |
| BaselineVisual Encoder Model=AIMv22025.12 | 86.02 | |
| Qwen3-VLModel Size=8B2026.03 | 85.7 | |
| Qwen2.5-VL-32BTraining=GSPO2026.03 | 85.7 | |
| LLaVA-OneVision-72BModel Scale=72B2024.08 | 85.6 | |
| LLaVA-OneVision-72B2025.06 | 85.6 | |
| Qwen2.5-VL-7BTraining=DF-GSPO2026.03 | 85.6 | |
| Latent DenoisingArchitecture=Qwen-2.5-VL2026.04 | 85.4 | |
| NVLM-72B2025.06 | 85.2 | |
| InternVL3-8BModel Category=Open-source General Models2025.11 | 85.2 | |
| InternVL3Params=8B2025.11 | 85.2 | |
| Qwen3-VL-30B-A3B-InstructInference scheme=no-thinking2026.04 | 85 | |
| VST-7B-SFTModel Category=Open-source SI Models2025.11 | 84.9 | |
| Qwen3-VLInference mode=Thinking mode, Size=8B2026.04 | 84.9 | |
| Qwen3VLParameters=8B, Internal Reasoning (Think mode)=true2026.05 | 84.9 | |
| GPT-4o (0806)Model Version=08062024.12 | 84.7 | |
| GPT-4o-05132024.09 | 84.6 | |
| GPT4oAccess=Close-source2026.02 | 84.6 | |
| GPT-4o (0513)2025.06 | 84.6 | |
| Qwen2.5-VL-32BTraining=Base2026.03 | 84.6 | |
| R1-ShareVL-7B2026.03 | 84.5 | |
| Qwen2.5-VL-7BTraining=GSPO2026.03 | 84.5 | |
| SenseNova-SIBase Architecture=Qwen3-VL-8B, Model Category=Ours2025.11 | 84.2 | |
| MiMo-EmbodiedParams=7B2025.11 | 84.2 | |
| Qwen3-VL + DiGLLM=Qwen3-8B2025.12 | 84.1 | |
| InternVL-3.5Model Size=8B2026.03 | 84 | |
| Qwen2.5-VL 72BLLM=Qwen2.5-72B2025.12 | 83.9 | |
| Qwen2.5-VL-7BTraining=Base2026.03 | 83.9 | |
| BaselineArchitecture=Qwen-2.5-VL2026.04 | 83.9 | |
| Qwen2.5-VLParams=7B2025.11 | 83.9 | |
| InternVL2Parameters=8B2026.03 | 83.8 | |
| VanillaBackbone=Qwen3-VL-8B, TFLOPs=100%2026.06 | 83.78 | |
| InternVL2-8B2024.09 | 83.6 | |
| Molmo-72BSize=72B2024.12 | 83.4 | |
| Molmo-72B2025.06 | 83.4 | |
| PerceptioParameters=8B2026.03 | 83.4 | |
| LaVerVisual Encoder Model=CLIP2025.12 | 83.3 | |
| SAYO-Qwen-8BBase Model=Qwen3-VL-8B2026.02 | 83.06 | |
| ThinkLite-7B2026.03 | 83 | |
| Qwen2.5-VL-7BBase Model=Qwen2.5-VL-7B2026.06 | 82.7 | |
| InternVL3.5-8BRatio=100%, Backbone=InternVL3.5-8B2026.05 | 82.67 | |
| OpenVLThinker-7BParameters=7B2026.02 | 82.61 | |
| GPT-4oParams=–2025.11 | 82.6 | |
| Op-SkipBackbone=Qwen3-VL-8B, TFLOPs=66%, N=202026.06 | 82.55 | |
| Ovis1.5-LLAMA3-8B2024.09 | 82.5 | |
| Qwen3-VLLLM=Qwen3-8B2025.12 | 82.5 | |
| VanillaBackbone=Qwen2.5-VL-7B, TFLOPs=100%2026.06 | 82.42 | |
| OneVision2024.09 | 82.4 | |
| GPT-4oModel Category=Proprietary VLMs2024.06 | 82.2 | |
| Qwen3-VL + DiGLLM=Qwen3-4B2025.12 | 82.2 | |
| Qwen3-VL-8B-ThinkingStrategy=LongCoT2026.02 | 82.2 | |
| Xuanwu VL-2BModel Size=2B2026.03 | 82.19 | |
| APETBackbone=Qwen3-VL-8B, TFLOPs=66%2026.06 | 82.16 | |
| Sa2VAParameters=8B2026.03 | 82.1 | |
| MiMo-VLParams=7B2025.11 | 81.83 | |
| ShortVBackbone=Qwen3-VL-8B, TFLOPs=58%, N=202026.06 | 81.77 | |
| Semantic-back-7BParameters=7B2026.02 | 81.74 | |
| V2DropBackbone=Qwen3-VL-8B, TFLOPs=66%2026.06 | 81.64 | |
| BaselineVisual Encoder Model=CLIP2025.12 | 81.61 | |
| IXC-2.52024.09 | 81.6 |