Table Question Answering on AIT-QA
93.5AccuracyOracle
Evaluation Results
| Method | Links | |
|---|---|---|
| OracleDescription=Theoretical upper bound via ideal selection between Textual and Symbolic reasoning2026.04 | 93.5 | |
| Formula-R1-PlusBackbone=Qwen 2.5-Coder 7B, Evaluation Protocol=Finetuning-Based, Distribution Setting=Out-of-distribution2025.05 | 93.2 | |
| ASTRA (Adaptive Selection)Reasoning mode=Adaptive2026.04 | 91.6 | |
| GraphOTTERCategory=Intermediate-Representation Methods2026.04 | 90.4 | |
| o3Category=Foundation Models2026.04 | 89.1 | |
| ASTRA (Symbolic Reasoning)Reasoning mode=Symbolic2026.04 | 87.3 | |
| E5Category=Prompting/Tool-augmented Methods2026.04 | 87.1 | |
| ASTRA (Textual Reasoning)Reasoning mode=Textual2026.04 | 86.1 | |
| TableGPT2IFT Data=2.36M, Backbone=Qwen2.5-7B2025.06 | 85.71 | |
| EEDPCategory=Prompting/Tool-augmented Methods2026.04 | 85.6 | |
| TableDreamerIFT Data=27K, Backbone=GPT-4o2025.06 | 84.29 | |
| MagpieIFT Data=100K2025.06 | 83.02 | |
| TableDreamerIFT Data=27K, Backbone=Llama3.1-70B-Instruct2025.06 | 82.99 | |
| GPT-4oCategory=Foundation Models2026.04 | 80.6 | |
| Formula-R1Backbone=Qwen 2.5-Coder 7B, Evaluation Protocol=Finetuning-Based, Distribution Setting=Out-of-distribution2025.05 | 80.39 | |
| Self-InstructIFT Data=100K2025.06 | 80.27 | |
| TableLLM-syn-dataIFT Data=80K2025.06 | 79.08 | |
| DeepSeek-V3Category=Foundation Models2026.04 | 78.5 | |
| Llama3.1-8B-Instruct2025.06 | 75.31 | |
| Gemini-3-Pro-PreviewReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 74.19 | |
| Evol-InstructIFT Data=100K2025.06 | 73.27 | |
| Phi3.5-mini-3.8B2025.06 | 72.98 | |
| Claude-Opus-4.5Reasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 72.22 | |
| Qwen3-32B-InstructReasoning Mode=thinking, Model Category=Reasoning Models2026.01 | 72.04 | |
| Doubao-1.5-thinking-proReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 71.19 | |
| Qwen3-8B-Think-SFT-TabCodeRLReasoning Mode=thinking, Model Category=Table-Specific Models2026.01 | 71.06 | |
| GPT-4oReasoning Mode=No reasoning, Model Category=No reasoning Models2026.01 | 70.53 | |
| GenQAIFT Data=100K2025.06 | 70.35 | |
| Deepseek-R1Reasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 69.05 | |
| OpenAI o1-miniReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 66.89 | |
| QWQ-32BReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 66.41 | |
| Qwen3-8B-InstructReasoning Mode=thinking, Model Category=Reasoning Models2026.01 | 66.21 | |
| TableLLMBackbone=Qwen 7B, Evaluation Protocol=Finetuning-Based2025.05 | 64.85 | |
| DeepSeek-V2-Lite-16B-Chat2025.06 | 64.72 | |
| GLM4-9B-Chat2025.06 | 63.45 | |
| ST-RaptorCategory=Intermediate-Representation Methods2026.04 | 62.7 | |
| TabAFBackbone=Qwen 2.5-Coder 7B, Evaluation Protocol=Finetuning-Based, Distribution Setting=Out-of-distribution2025.05 | 62.33 | |
| Qwen3-8B-InstructReasoning Mode=no-thinking, Model Category=No reasoning Models2026.01 | 61.55 | |
| DynasourIFT Data=132K2025.06 | 59.66 | |
| Qwen2.5-72B-InstructReasoning Mode=No reasoning, Model Category=No reasoning Models2026.01 | 58.59 | |
| MiniCPM3-4B2025.06 | 57.82 | |
| InternLM2.5-7B-Chat2025.06 | 55.97 | |
| Yi-1.5-9B-Chat2025.06 | 53.87 | |
| OmniTabIFT Data=100K2025.06 | 50.63 | |
| TableGPT2-7BReasoning Mode=Table-Specific, Model Category=Table-Specific Models2026.01 | 49.89 | |
| ReasTapIFT Data=100K2025.06 | 48.49 | |
| TableGPT-syn-dataIFT Data=66K2025.06 | 47.52 | |
| Llama-3.1-8B-InstructReasoning Mode=No reasoning, Model Category=No reasoning Models2026.01 | 44.8 | |
| Baichuan2-7B-Chat2025.06 | 41.59 | |
| Mistral-7B-Instruct-v0.32025.06 | 36.25 | |
| UCTRIFT Data=43K2025.06 | 35.76 | |
| TableBenchLLMIFT Data=20K, Backbone=Llama3.1-8B2025.06 | 30.41 | |
| TableLlamaBackbone=Llama-2 7B, Evaluation Protocol=Finetuning-Based, Distribution Setting=Out-of-distribution2025.05 | 26.99 | |
| TableLLMIFT Data=80K, Backbone=CodeLlama-7B2025.06 | 24.87 | |
| TableLLM-7BReasoning Mode=Table-Specific, Model Category=Table-Specific Models2026.01 | 21.32 | |
| Deepseek-R1-Distill-Qwen-7BReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 18.06 | |
| TableGPT2Backbone=Qwen 2.5 7B, Evaluation Protocol=Finetuning-Based, Distribution Setting=Out-of-distribution2025.05 | 12.43 | |
| TableLlama-7BReasoning Mode=Table-Specific, Model Category=Table-Specific Models2026.01 | 11.33 |