Table Question Answering on HiTab
94.1AccuracyOracle
Evaluation Results
| Method | Links | |
|---|---|---|
| OracleDescription=Theoretical upper bound via ideal selection between Textual and Symbolic reasoning2026.04 | 94.1 | |
| ASTRA (Adaptive Selection)Reasoning mode=Adaptive2026.04 | 90.1 | |
| ASTRA (Symbolic Reasoning)Reasoning mode=Symbolic2026.04 | 89.3 | |
| GraphOTTERCategory=Intermediate-Representation Methods2026.04 | 88.8 | |
| Gemini-3-Pro-PreviewReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 88.35 | |
| Doubao-1.5-thinking-proReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 87.28 | |
| OpenAI o1-miniReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 87.11 | |
| Claude-Opus-4.5Reasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 86.45 | |
| Qwen3-32B-InstructReasoning Mode=thinking, Model Category=Reasoning Models2026.01 | 85.57 | |
| o3Category=Foundation Models2026.04 | 85.3 | |
| E5Category=Prompting/Tool-augmented Methods2026.04 | 85.1 | |
| GPT-4oReasoning Mode=No reasoning, Model Category=No reasoning Models2026.01 | 83.55 | |
| Deepseek-R1Reasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 82.99 | |
| ASTRA (Textual Reasoning)Reasoning mode=Textual2026.04 | 82.2 | |
| QWQ-32BReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 82.11 | |
| DeepSeek-V3Category=Foundation Models2026.04 | 82 | |
| Qwen3-VL-8B + TABQAWORLDStrategy=Training-free Agent2026.04 | 81.41 | |
| Qwen2.5-72B-InstructReasoning Mode=No reasoning, Model Category=No reasoning Models2026.01 | 81.03 | |
| Qwen3-8B-Think-SFT-TabCodeRLReasoning Mode=thinking, Model Category=Table-Specific Models2026.01 | 80.72 | |
| EEDPCategory=Prompting/Tool-augmented Methods2026.04 | 79.2 | |
| GPT-4oCategory=Foundation Models2026.04 | 78.6 | |
| E5Table Representation=Original Hierarchical Tables (Direct)2025.01 | 77.3 | |
| Qwen2.5-VL-7B + TABQAWORLDStrategy=Training-free Agent2026.04 | 76.88 | |
| GRABSetting=Tuned LLM (LoRA)2026.06 | 76.83 | |
| zero-shotSetting=Tuned LLM (LoRA)2026.06 | 76.01 | |
| TableGPT2-72BCategory=Table-Specific Adapted Models, Model size=72B2026.04 | 75.6 | |
| Qwen3-8B-InstructReasoning Mode=thinking, Model Category=Reasoning Models2026.01 | 75.37 | |
| GRABSetting=Frozen LLM2026.06 | 74.49 | |
| TableDART (TG2-7B+Ovis2-8B)Strategy=Dynamic Adaptive Routing2026.04 | 74.37 | |
| TableMasterTable Representation=Converted Relational Tables, Backbone=GPT-4o2025.01 | 74.2 | |
| TableGPT2IFT Data=2.36M, Backbone=Qwen2.5-7B2025.06 | 73.97 | |
| TAMOSetting=Frozen LLM, Backbone=Llama 3.1 8B2025.11 | 73.73 | |
| MultiCoT (optimized prompt + verbalized table)Table Representation=Converted Relational Tables, Prompting Strategy=optimized prompt, Verbalization=verbalized table, Backbone=GPT-4o2025.01 | 73.5 | |
| GPT-5.4-miniSetting=Others2026.06 | 72.28 | |
| VD-LF (HQ vs. LQ)Model=Qwen2.5-7B-VL, Train Set=HiTab2026.03 | 71.91 | |
| VD-LB (HQ Correct vs. LQ Wrong)Model=Qwen2.5-7B-VL, Train Set=HiTab2026.03 | 71.4 | |
| TableDART (TG2-7B+Qwen2.5-VL-7B)Strategy=Dynamic Adaptive Routing2026.04 | 71.13 | |
| VD-LB (HQ Correct vs. LQ Wrong)Model=Qwen2.5-7B-VL, Train Set=VQA2026.03 | 70.58 | |
| TableGPT2-7B (Text-only Path)Strategy=Dynamic Adaptive Routing2026.04 | 70.27 | |
| VD-LF (HQ vs. LQ)Model=Qwen2.5-7B-VL, Train Set=GQA2026.03 | 70.2 | |
| MultiCoT (optimized prompt)Table Representation=Converted Relational Tables, Prompting Strategy=optimized prompt, Backbone=GPT-4o2025.01 | 70 | |
| SFT (HQ Correct Only)Model=Qwen2.5-7B-VL, Train Set=HiTab2026.03 | 69.89 | |
| Prompt tuningSetting=Frozen LLM2026.06 | 69.63 | |
| Prompt tuningSetting=Frozen LLM, Backbone=Llama 3.1 8B2025.11 | 69.38 | |
| VD-LB (HQ Correct vs. LQ Wrong)Model=Qwen2.5-7B-VL, Train Set=GQA2026.03 | 69.19 | |
| VD-LF (HQ vs. LQ)Model=Qwen2.5-7B-VL, Train Set=VQA2026.03 | 68.81 | |
| Ovis2-8B (Image-only Path)Strategy=Dynamic Adaptive Routing2026.04 | 68.59 | |
| BaselineModel=Qwen2.5-7B-VL, Training=Inference Only2026.03 | 67.93 | |
| Qwen3-VL-8BModality=Image2026.04 | 65.83 | |
| SFT (HQ Correct Only)Model=Qwen2.5-7B-VL, Train Set=VQA2026.03 | 65.4 | |
| TAMOSetting=Frozen LLM2026.06 | 65.27 | |
| TableLlama+textLLM=Llama-2 7B, Representation=OCR textual table2026.02 | 64.71 | |
| Specialist SOTASetting=Others2025.11 | 64.71 | |
| TableLlamaCategory=Table-Specific Adapted Models2026.04 | 64.7 | |
| SFT (HQ Correct Only)Model=Qwen2.5-7B-VL, Train Set=GQA2026.03 | 64.27 | |
| MultiCoT (original)Table Representation=Converted Relational Tables, Prompting Strategy=original, Backbone=GPT-4o2025.01 | 64 | |
| TAMO+Setting=Tuned LLM (SFT)2025.11 | 63.89 | |
| TableLLamaSetting=Others2026.06 | 63.82 | |
| TableLlamaSetting=Tuned LLM (SFT)2025.11 | 63.76 | |
| SFT (HQ Correct Only)Model=Qwen2.5-3B-VL, Train Set=HiTab2026.03 | 63.45 | |
| HIPPO-8BModality=Multimodal2026.04 | 63 | |
| Qwen2.5-VL-7BModality=Image2026.04 | 62.69 | |
| Qwen3-8B-InstructReasoning Mode=no-thinking, Model Category=No reasoning Models2026.01 | 61.62 | |
| VD-LF (HQ vs. LQ)Model=Qwen2.5-3B-VL, Train Set=HiTab2026.03 | 61.62 | |
| VD-LB (HQ Correct vs. LQ Wrong)Model=Qwen2.5-3B-VL, Train Set=GQA2026.03 | 61.36 | |
| VD-LB (HQ Correct vs. LQ Wrong)Model=Qwen2.5-3B-VL, Train Set=HiTab2026.03 | 61.11 | |
| GPT-4.1Setting=Others2025.11 | 60.54 | |
| Gemini 2.0 FlashModality=Multimodal2026.04 | 60.41 | |
| VD-LF (HQ vs. LQ)Model=Qwen2.5-3B-VL, Train Set=VQA2026.03 | 60.23 | |
| VD-LF (HQ vs. LQ)Model=Qwen2.5-3B-VL, Train Set=GQA2026.03 | 60.23 | |
| TAMO+Setting=Tuned LLM (LoRA)2025.11 | 59.22 | |
| VD-LB (HQ Correct vs. LQ Wrong)Model=Qwen2.5-3B-VL, Train Set=VQA2026.03 | 59.09 | |
| BaselineModel=Qwen2.5-3B-VL, Training=Inference Only2026.03 | 58.96 | |
| Llama-3.1-8B-InstructReasoning Mode=No reasoning, Model Category=No reasoning Models2026.01 | 57.21 | |
| GenQAIFT Data=100K2025.06 | 57.14 | |
| TableDreamerIFT Data=27K, Backbone=Llama3.1-70B-Instruct2025.06 | 56.75 | |
| MiniCPM-V-2.6-8BModality=Image2026.04 | 56.53 | |
| SFT (HQ Correct Only)Model=Qwen2.5-3B-VL, Train Set=VQA2026.03 | 56.25 | |
| SFT (HQ Correct Only)Model=Qwen2.5-3B-VL, Train Set=GQA2026.03 | 55.62 | |
| MiniCPM3-4B2025.06 | 55.34 | |
| SFTSetting=Tuned LLM (SFT)2025.11 | 54.8 | |
| zero-shotSetting=Inference-only2026.06 | 53.72 | |
| TableDreamerIFT Data=27K, Backbone=GPT-4o2025.06 | 53.22 | |
| Mistral-7B-Instruct-v0.32025.06 | 52.05 | |
| Yi-1.5-9B-Chat2025.06 | 51.85 | |
| InternLM2.5-7B-Chat2025.06 | 51.46 | |
| LoRASetting=Tuned LLM (LoRA)2025.11 | 50.76 | |
| ST-RaptorCategory=Intermediate-Representation Methods2026.04 | 49 | |
| Self-InstructIFT Data=100K2025.06 | 48.92 | |
| TableGPT2-7BReasoning Mode=Table-Specific, Model Category=Table-Specific Models2026.01 | 48.87 | |
| TAMOSetting=Frozen LLM2025.11 | 48.86 | |
| GPT-4Setting=Others2025.11 | 48.4 | |
| Deepseek-R1-Distill-Qwen-7BReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 47.78 | |
| MagpieIFT Data=100K2025.06 | 47.16 | |
| TableLlama-7BModality=Text2026.04 | 46.57 | |
| TableLLMIFT Data=80K, Backbone=CodeLlama-7B2025.06 | 45.4 | |
| Evol-InstructIFT Data=100K2025.06 | 45.2 | |
| Llama3.1-8B-Instruct2025.06 | 43.63 | |
| GPT-3.5Setting=Others2025.11 | 43.62 | |
| DynasourIFT Data=132K2025.06 | 43.44 |