Table Question Answering on ReachQA (test)
74.85Relaxed AccuracyHuman
Evaluation Results
| Method | Links | |
|---|---|---|
| HumanModel Category=Baseline, Evaluation Source=Model authors or leaderboards2025.09 | 74.85 | |
| Claude 3.5 SonnetModel Category=Proprietary VLMs, Evaluation Source=Model authors or leaderboards2025.09 | 63 | |
| Gemini 2.5 ProModel Category=Proprietary VLMs, Evaluation Source=LLM jury2025.09 | 61.87 | |
| Qwen2.5-VL-7B-Instruct + Visual-TableQAModel Category=Finetuned VLMs, Evaluation Source=LLM jury2025.09 | 60.95 | |
| Gemini 2.5 FlashModel Category=Proprietary VLMs, Evaluation Source=LLM jury2025.09 | 56.97 | |
| Qwen2.5-VL-7B-Instruct + ReachQAModel Category=Finetuned VLMs, Evaluation Source=LLM jury2025.09 | 55.75 | |
| GPT-4oModel Category=Proprietary VLMs, Evaluation Source=Model authors or leaderboards2025.09 | 53.25 | |
| Qwen2.5-VL-32B-InstructModel Category=Open-Source VLMs, Evaluation Source=LLM jury2025.09 | 49.5 | |
| Qwen2.5-VL-7B-InstructModel Category=Open-Source VLMs, Evaluation Source=LLM jury2025.09 | 49.23 | |
| Llama 4 Maverick 17B-128E InstructModel Category=Open-Source VLMs, Evaluation Source=LLM jury2025.09 | 47.98 | |
| Mistral Small 3.1 24B InstructModel Category=Open-Source VLMs, Evaluation Source=Model authors or leaderboards2025.09 | 42.45 | |
| GPT-4o miniModel Category=Proprietary VLMs, Evaluation Source=Model authors or leaderboards2025.09 | 40.35 |