Table Question Answering on WikiTableQuestions (test)
76.6AccuracyTabLaP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TabLaP2025.03 | 76.6 | — | — | |
| RePandamode=Finetuned-Pandas for QA, training examples=1,2002025.03 | 75.1 | — | — | |
| SynTQAengine=GPT2025.03 | 74.4 | — | — | |
| Mix SC2025.03 | 73.6 | — | — | |
| SynTQAengine=RF2025.03 | 71.6 | — | — | |
| CABINET2025.03 | 69.1 | — | — | |
| GPT-42025.01 | 68.4 | — | — | |
| Chain-of-Table2025.03 | 67.31 | — | — | |
| Tab-PoT2025.03 | 66.78 | — | — | |
| Codex w/ DATEREvaluation Protocol=LLM based2023.01 | 65.9 | — | — | |
| OmniTabf-shot=full, pretraining_data=all2022.07 | 62.8 | — | — | |
| OmniTab w/ DATEREvaluation Protocol=Fine-tuning2023.01 | 62.5 | — | — | |
| BinderEvaluation Protocol=LLM based2023.01 | 61.9 | — | — | |
| OmniTabf-shot=full, pretraining_data=natural2022.07 | 61.3 | — | — | |
| OmniTabf-shot=full, pretraining_data=synthetic2022.07 | 61.3 | — | — | |
| OmniTabEvaluation Protocol=Fine-tuning2023.01 | 61.2 | — | — | |
| TaCubeEvaluation Protocol=Fine-tuning2023.01 | 60.8 | — | — | |
| TAPEX*f-shot=full2022.07 | 60.1 | — | — | |
| TAPEXf-shot=full2022.07 | 59.5 | — | — | |
| ReasTAPEvaluation Protocol=Fine-tuning2023.01 | 58.6 | — | — | |
| TAPEXEvaluation Protocol=Fine-tuning2023.01 | 57.2 | — | — | |
| TableRAG2024.10 | 57.03 | — | — | |
| Binder2024.10 | 56.74 | — | — | |
| GPT-3.52025.01 | 53.13 | — | — | |
| Text-to-SQL2024.10 | 52.9 | — | — | |
| TAMAbackbone=LLaMA 3.1 8B Instruct2025.01 | 52.88 | — | — | |
| Dater2024.10 | 52.81 | — | — | |
| Wang et al. (2019) + GRAPPA (MLM+SSP)Encoder=GRAPPA, Pre-training objective=MLM+SSP2020.09 | 52.7 | — | — | |
| TableFormerEvaluation Protocol=Fine-tuning2023.01 | 52.6 | — | — | |
| TaBERT+MAPOf-shot=full2022.07 | 52.3 | — | — | |
| TaBERT2024.10 | 52.3 | — | — | |
| OmniTabf-shot=1024, pretraining_data=all2022.07 | 51.9 | — | — | |
| Yin et al. (2020b)2020.09 | 51.8 | — | — | |
| Wang et al. (2019) + GRAPPA (MLM)Encoder=GRAPPA, Pre-training objective=MLM2020.09 | 51.7 | — | — | |
| Wang et al. (2019) + GRAPPA (SSP)Encoder=GRAPPA, Pre-training objective=SSP2020.09 | 51.1 | — | — | |
| Wang et al. (2019) + RoBERTa-largeEncoder=RoBERTa-large2020.09 | 50.9 | — | — | |
| TaPasEvaluation Protocol=Fine-tuning2023.01 | 50.4 | — | — | |
| OmniTabf-shot=1024, pretraining_data=natural2022.07 | 49.8 | — | — | |
| T5-3BEvaluation Protocol=Fine-tuning2023.01 | 49.3 | — | — | |
| Herzig et al. (2020b)2020.09 | 48.8 | — | — | |
| TAPASf-shot=full2022.07 | 48.8 | — | — | |
| OmniTabf-shot=1024, pretraining_data=synthetic2022.07 | 48.8 | — | — | |
| CodexEvaluation Protocol=LLM based2023.01 | 47.6 | — | — | |
| Structured-Attention ParserEnsemble size=102019.09 | 47.3 | — | — | |
| Agarwal et al. (2019)Ensemble size=102019.09 | 46.9 | — | — | |
| Liang et al. (2018)Ensemble size=102019.09 | 46.3 | — | — | |
| TAPEXf-shot=10242022.07 | 45.5 | — | — | |
| IterativeSearchEvaluation Protocol=Fine-tuning2023.01 | 44.7 | — | — | |
| TAPEX*f-shot=10242022.07 | 44.6 | — | — | |
| Wang et al. (2019)2020.09 | 44.5 | — | — | |
| LatentAlignmentEvaluation Protocol=Fine-tuning2023.01 | 44.5 | — | — | |
| Agarwal et al. (2019)2020.09 | 44.1 | — | — | |
| MeRLEvaluation Protocol=Fine-tuning2023.01 | 44.1 | — | — | |
| Dasigi et al. (2019)2020.09 | 43.9 | — | — | |
| MAPOEvaluation Protocol=Fine-tuning2023.01 | 43.8 | — | — | |
| LLaMA 3.1 8B Instructstatus=base model2025.01 | 43.46 | — | — | |
| Liang et al. (2018)2020.09 | 43.1 | — | — | |
| Wang et al. (2019) + GRAPPA (MLM+SSP) (10% data)Encoder=GRAPPA, Pre-training objective=MLM+SSP, Training data percentage=10%2020.09 | 42 | — | — | |
| OmniTabf-shot=128, pretraining_data=all2022.07 | 41.4 | — | — | |
| OmniTabf-shot=128, pretraining_data=natural2022.07 | 38.4 | — | — | |
| Wang et al. (2019) + RoBERTa-large (10% data)Encoder=RoBERTa-large, Training data percentage=10%2020.09 | 38.1 | — | — | |
| BARTf-shot=full2022.07 | 38 | — | — | |
| OmniTabf-shot=128, pretraining_data=synthetic2022.07 | 37.5 | — | — | |
| mPLUG-Owl-7B + OursModel Params=7.2B, Trainable Params=96M, Pre-training Data=1.1B, Fine-tuning Data=650K2024.11 | 34.5 | — | — | |
| TAPASf-shot=10242022.07 | 33.6 | — | — | |
| TaBERT+MAPOf-shot=10242022.07 | 33.3 | — | — | |
| MonkeyModel Params=9.8B, Trainable Params=207M, Pre-training Data=1.4B+76.8M, Fine-tuning Data=1.44M2024.11 | 32.8 | — | — | |
| mPLUG-Owl-7B + UReaderModel Params=7.2B, Trainable Params=86M, Pre-training Data=1.1B, Fine-tuning Data=650K2024.11 | 29.4 | — | — | |
| OmniTabf-shot=16, pretraining_data=all2022.07 | 26.8 | — | — | |
| TAPEX*f-shot=1282022.07 | 25.2 | — | — | |
| TAPEXf-shot=1282022.07 | 23.1 | — | — | |
| OmniTabf-shot=16, pretraining_data=natural2022.07 | 22.8 | — | — | |
| OmniTabf-shot=16, pretraining_data=synthetic2022.07 | 21.5 | — | — | |
| BLIP-2-OPT-2.7B + OursModel Params=3.8B, Trainable Params=14M, Pre-training Data=129M, Fine-tuning Data=650K2024.11 | 21.2 | — | — | |
| TAPASf-shot=1282022.07 | 18.9 | — | — | |
| DonutModel Params=143M, Trainable Params=143M, Pre-training Data=13M, Fine-tuning Data=Task-specific2024.11 | 18.8 | — | — | |
| BLIP-2-OPT-2.7B + UReaderModel Params=3.8B, Trainable Params=8M, Pre-training Data=129M, Fine-tuning Data=650K2024.11 | 17.4 | — | — | |
| BARTf-shot=10242022.07 | 17.3 | — | — | |
| Qwen-VLModel Params=9.6B, Trainable Params=207M, Pre-training Data=1.4B+76.8M, Fine-tuning Data=Task-specific2024.11 | 15.9 | — | — | |
| TAPEX*f-shot=162022.07 | 15.7 | — | — | |
| TaBERT+MAPOf-shot=1282022.07 | 15.1 | — | — | |
| TAPEXf-shot=162022.07 | 10.4 | — | — | |
| TAPASf-shot=162022.07 | 9.8 | — | — | |
| BARTf-shot=1282022.07 | 8.4 | — | — | |
| TaBERT+MAPOf-shot=162022.07 | 7.7 | — | — | |
| BARTf-shot=162022.07 | 2.9 | — | — | |
| Agarwal et al. (2019)Category=Previous Systems2021.07 | — | 44.1 | — | |
| Agarwal et al. (2019)2020.04 | — | 44.1 | — | |
| BARTCategory=Pre-trained Language Models2021.07 | — | 38 | — | |
| BinderLLM=Codex2024.06 | — | — | 61.9 | |
| CoTLLM=Mixtral-8x7B2024.06 | — | — | 53.48 | |
| CoTLLM=Mistral-7B2024.06 | — | — | 30.46 | |
| CoTLLM=DeepSeek-67B2024.06 | — | — | 55.57 | |
| CoTLLM=DeepSeek-7B2024.06 | — | — | 34.65 | |
| Dasigi et al. (2019)Category=Previous Systems2021.07 | — | 44.3 | — | |
| Dasigi et al. (2019)2020.04 | — | 43.9 | — | |
| DaterLLM=Codex2024.06 | — | — | 65.9 | |
| DirectLLM=Mixtral-8x7B2024.06 | — | — | 53.08 | |
| DirectLLM=Mistral-7B2024.06 | — | — | 27.19 | |
| DirectLLM=DeepSeek-67B2024.06 | — | — | 54.72 |