Visual Question Answering on AI2D
91.81AccuracyGR3D-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| GR3D-8BTraining Stage=Stage 12026.05 | 91.81 | |
| GR3D-8BTraining Stage=Stage 22026.05 | 91.54 | |
| NVILA-Lite-8B2026.05 | 91.01 | |
| Gemini 2.5 Pro2026.04 | 88.4 | |
| OpenVLThinkerV2Parameters=8B2026.04 | 87.5 | |
| BaselineModel=Llama Nemotron Nano 12B v2 VL2026.01 | 87.3 | |
| VISEModel Scale=32B2026.06 | 87.24 | |
| VisionZero-ChartModel Scale=32B2026.06 | 87.16 | |
| VisionZero-CLEVRModel Scale=32B2026.06 | 87.1 | |
| iReasonerModel Scale=32B2026.06 | 87.08 | |
| EvoLMMModel Scale=32B2026.06 | 87.02 | |
| PTQModel=Llama Nemotron Nano 12B v2 VL2026.01 | 86.8 | |
| VisPlayModel Scale=32B2026.06 | 86.71 | |
| QADModel=Llama Nemotron Nano 12B v2 VL2026.01 | 86.7 | |
| VisionZero-RWModel Scale=32B2026.06 | 86.69 | |
| BaseModel Scale=32B2026.06 | 86.63 | |
| QATModel=Llama Nemotron Nano 12B v2 VL2026.01 | 86.5 | |
| Qwen3-VL-8BRetained Tokens=1024 Tokens (100%)2026.07 | 86.3 | |
| EMOVAModel Size=72B2024.09 | 85.8 | |
| InternVL3Input=Tile-wise, RoPE=1D-RoPE, Backbone=Dense, Parameter Scale=8B2025.10 | 85.2 | |
| OneThinker-8BParameters=8B2026.04 | 85.2 | |
| VanillaSource=Qwen’25, Backbone=Qwen3.5-9B-Hybrid, Average Gated Attention Token Reduction=0%2026.06 | 85.2 | |
| Qwen3-VL GRPORL Strategy=GRPO2026.04 | 85.1 | |
| Qwen3-VL GDPORL Strategy=GDPO2026.04 | 85.1 | |
| GPT-4o2026.04 | 84.9 | |
| VisionZero2026.04 | 84.8 | |
| GPT-4o2024.09 | 84.6 | |
| InternVL2.5Input=Tile-wise, RoPE=1D-RoPE, Backbone=Dense, Parameter Scale=8B2025.10 | 84.5 | |
| MM-Eureka-7BParameters=7B2026.04 | 84.1 | |
| VISEModel Scale=8B2026.06 | 84.1 | |
| VisionZero-ChartModel Scale=8B2026.06 | 84.04 | |
| Qwen3-VL-4BModel=Qwen3-VL, Parameter Count=4B2026.04 | 84 | |
| VisionZero-CLEVRModel Scale=8B2026.06 | 83.98 | |
| iReasonerModel Scale=8B2026.06 | 83.96 | |
| Qwen2.5-VLInput=Any Res., RoPE=M-RoPE, Backbone=Dense, Parameter Scale=8B2025.10 | 83.9 | |
| EvoLMMModel Scale=8B2026.06 | 83.89 | |
| PTPModel=InternVL2-8B, Pruning Method=PTP, Pruning Ratio (r)=0.5, Balancing Weight (alpha)=0.52025.09 | 83.8 | |
| FLARE-X 8B# Vis tok.=14002025.04 | 83.6 | |
| VisPlayModel Scale=8B2026.06 | 83.45 | |
| Qwen2.5VL 7B# Vis tok.=14002025.04 | 83.4 | |
| VisionZero-RWModel Scale=8B2026.06 | 83.37 | |
| BaseModel Scale=8B2026.06 | 83.31 | |
| NEOInput=Any Res., RoPE=Native-RoPE, Backbone=Dense, Parameter Scale=8B2025.10 | 83.1 | |
| MoVALLM=Hermes-Yi-34B, Params=38B, Model Type=Specialist2024.04 | 83 | |
| Qwen2-VLInput=Any Res., RoPE=M-RoPE, Backbone=Dense, Parameter Scale=8B2025.10 | 83 | |
| Encoder-BasedInput=Tile-wise, RoPE=1D-RoPE, Backbone=Dense, Parameter Scale=8B2025.10 | 82.9 | |
| VanillaSource=Qwen’25, Token Reduction=0%, Backbone=Qwen2.5-VL-7B2026.06 | 82.6 | |
| LLaVA-OneVision-7BModel=LLaVA-OneVision, Parameter Count=7B2026.04 | 82.51 | |
| FP16 BaselineModel=InternVL2-8B, Bitwidth=FP162026.03 | 82.42 | |
| InternVL2-8BModel=InternVL2-8B, Pruning Method=None, Pruning Ratio (r)=1.02025.09 | 82.4 | |
| RADIOv2.5-H + TilingVision Encoder=RADIOv2.5-H (ours) + Tiling, Resolution=Up to 7 × 768², Compression=ToMe r=2108, Tokens/im=~ 12332024.12 | 82.4 | |
| Qwen3-VL-Instruct-8BParameters=8B2026.04 | 82.3 | |
| VISEModel Scale=4B2026.06 | 82.16 | |
| GSearchModel=InternVL2-8B, Pruning Method=GSearch2025.09 | 82.1 | |
| Vision-G12026.04 | 82.1 | |
| VTWModel=InternVL2-8B, Pruning Method=VTW2025.09 | 81.8 | |
| OpenVLThinker-7BParameters=7B2026.04 | 81.8 | |
| VisionZero-ChartModel Scale=4B2026.06 | 81.76 | |
| EMOVAModel Size=7B2024.09 | 81.7 | |
| Qwen2.5-VL 3B (reported)Res.=Native2025.12 | 81.6 | |
| Qwen2.5-VLInput=Any Res., RoPE=M-RoPE, Backbone=Dense, Parameter Scale=2B2025.10 | 81.6 | |
| Qwen3-VL-4B-Instruct2026.02 | 81.6 | |
| LLaVA-OneVision 7B# Vis tok.=24002025.04 | 81.6 | |
| FLARE-L 8B# Vis tok.=6302025.04 | 81.6 | |
| GeminiProVision2024.06 | 81.4 | |
| Qwen2.5VL 3B# Vis tok.=14002025.04 | 81.4 | |
| FP16 BaselineModel=LLaVA-onevision-7B, Bitwidth=FP162026.03 | 81.31 | |
| iReasonerModel Scale=4B2026.06 | 81.22 | |
| PALI-X-55BModel type=Specialist SOTAs, Configuration=Single-task fine-tuning, without OCR Pipeline, Parameters=55B2023.08 | 81.2 | |
| PALI-X-55BParams=55B, Model Type=Specialist2024.04 | 81.2 | |
| FLARE-X 3B# Vis tok.=14002025.04 | 81.2 | |
| EvoLMMModel Scale=4B2026.06 | 81.13 | |
| MG-LLAVALLM=Yi1.5-34B, Parameters=34.4B, Resolution=336(768)2024.06 | 81.1 | |
| FastVModel=InternVL2-8B, Pruning Method=FastV, Pruning Ratio (r)=0.52025.09 | 81.1 | |
| Molmo 7B# Vis tok.=12002025.04 | 81 | |
| VisionZero-CLEVRModel Scale=4B2026.06 | 80.81 | |
| VL-Rethinker-7BParameters=7B2026.04 | 80.8 | |
| SpB2.0-VL-5BModel=SpB2.0-VL, Parameter Count=5B2026.04 | 80.73 | |
| RTNModel=InternVL2-8B, Bitwidth=W3A162026.03 | 80.51 | |
| Qwen2.5-VL-3BModel=Qwen2.5-VL, Parameter Count=3B2026.04 | 80.51 | |
| PDropModel=InternVL2-8B, Pruning Method=PDrop2025.09 | 80.4 | |
| VisPlayModel Scale=4B2026.06 | 80.34 | |
| VisionZero-RWModel Scale=4B2026.06 | 80.31 | |
| Gemini Pro 1.52024.09 | 80.3 | |
| FP16 BaselineModel=Qwen2-VL-7B, Bitwidth=FP162026.03 | 80.12 | |
| Qwen2.5-VL 3B (reproduced)Res.=≤ 896²2025.12 | 80.1 | |
| NEOInput=Any Res., RoPE=Native-RoPE, Backbone=Dense, Parameter Scale=2B2025.10 | 80.1 | |
| BaseModel Scale=4B2026.06 | 80.1 | |
| QIGModel=InternVL2-8B, Bitwidth=W3A162026.03 | 79.73 | |
| MBQModel=InternVL2-8B, Bitwidth=W3A162026.03 | 79.66 | |
| QIGModel=InternVL2-8B, Bitwidth=W4A82026.03 | 79.63 | |
| AWQModel=InternVL2-8B, Bitwidth=W3A162026.03 | 79.47 | |
| MBQModel=InternVL2-8B, Bitwidth=W4A82026.03 | 79.47 | |
| FLARE-L 3B# Vis tok.=6302025.04 | 79.4 | |
| VITAModel Size=1.52024.09 | 79.3 | |
| SigLIP SO400M + TilingVision Encoder=SigLIP SO400M [52] + Tiling, Resolution=Up to 13 × 384², Compression=2 × 2 Unshuffle, Tokens/im=~ 19282024.12 | 79.3 | |
| Eagle2number of parameters=1.5B2025.07 | 79.3 | |
| QIGModel=LLaVA-onevision-7B, Bitwidth=W3A162026.03 | 79.11 | |
| InstructVLA-Generalistnumber of parameters=1.5B2025.07 | 79.1 | |
| RTNModel=InternVL2-8B, Bitwidth=W4A82026.03 | 79.02 |