Visual Question Answering on ChartQA (val)
81.9Accuracyjina-vlm
Evaluation Results
| Method | Links | |
|---|---|---|
| jina-vlm2025.12 | 81.9 | |
| InternVL3.5-2BParameter Count=2B2025.12 | 80.7 | |
| InternVL3-2BParameter Count=2B2025.12 | 80.2 | |
| Qwen3-VL-2BParameter Count=2B2025.12 | 77.2 | |
| Qwen2-VL-2BParameter Count=2B2025.12 | 73.5 | |
| DeepStack-HD+LLM=Vicuna-13B, Visual Tokens=14400, Context Length=2880, Pre-training=748K, Instruction-tuning=748K2024.06 | 64 | |
| LLaVA-NextLLM=Vicuna-13B, Visual Tokens=2880, Context Length=2880, Pre-training=558K, Instruction-tuning=765K2024.06 | 62.2 | |
| VisionTaPasData Table Setting=Gold Data Table Provided2022.03 | 59.32 | |
| T5Data Table Setting=Gold Data Table Provided2022.03 | 59.11 | |
| VL-T5Data Table Setting=Gold Data Table Provided2022.03 | 58.8 | |
| DeepStack-HD+LLM=Vicuna-7B, Visual Tokens=14400, Context Length=2880, Pre-training=558K, Instruction-tuning=748K2024.06 | 56.3 | |
| LLaVA-NextLLM=Vicuna-7B, Visual Tokens=2880, Context Length=2880, Pre-training=558K, Instruction-tuning=765K2024.06 | 54.8 | |
| TaPasData Table Setting=Gold Data Table Provided2022.03 | 49.16 | |
| VisionTaPasData Table Setting=Gold Data Table Not Provided2022.03 | 42.6 | |
| T5Data Table Setting=Gold Data Table Not Provided2022.03 | 40.15 | |
| TaPasData Table Setting=Gold Data Table Not Provided2022.03 | 39.68 | |
| VL-T5Data Table Setting=Gold Data Table Not Provided2022.03 | 38.43 | |
| PlotQA*Data Table Setting=Gold Data Table Not Provided2022.03 | 36.15 | |
| DeepStack-LLLM=Vicuna-13B, Visual Tokens=2880, Context Length=576, Pre-training=558K, Instruction-tuning=665K2024.06 | 21.2 | |
| DeepStack-LLLM=Vicuna-7B, Visual Tokens=2880, Context Length=576, Pre-training=558K, Instruction-tuning=665K2024.06 | 21 | |
| DeepStack-VLLM=Vicuna-7B, Visual Tokens=2880, Context Length=576, Pre-training=558K, Instruction-tuning=665K2024.06 | 20.6 | |
| DeepStack-VLLM=Vicuna-13B, Visual Tokens=2880, Context Length=576, Pre-training=558K, Instruction-tuning=665K2024.06 | 20.2 | |
| LLaVA-1.5LLM=Vicuna-7B, Visual Tokens=576, Context Length=576, Pre-training=558K, Instruction-tuning=665K2024.06 | 18.2 | |
| LLaVA-1.5LLM=Vicuna-13B, Visual Tokens=576, Context Length=576, Pre-training=558K, Instruction-tuning=665K2024.06 | 18.2 | |
| PREFILData Table Setting=Gold Data Table Not Provided2022.03 | 4.53 |