Table Fact Verification on TabFact (test)
93.73AccuracyTableMind++
Evaluation Results
| Method | Links | |
|---|---|---|
| TableMind++Model Category=Tuning-based, Pass@1=true2026.03 | 93.73 | |
| TableMindModel Category=Tuning-based, Pass@1=true2026.03 | 91.85 | |
| TableMasterBackbone=Llama-3.170B, Evaluation Protocol=Single inference run2025.01 | 91.16 | |
| TableMasterBackbone=gpt-4o-mini∼8B, Evaluation Protocol=Single inference run2025.01 | 90.12 | |
| GPT-5Model Category=Proprietary, Pass@1=true2026.03 | 90.05 | |
| Gemini-2.5-flashModel Category=Proprietary, Pass@1=true2026.03 | 89.65 | |
| PASTA2022.11 | 89.3 | |
| EnoTab2025.09 | 89.2 | |
| PoTableModel Category=Training-free, Pass@1=true2026.03 | 88.93 | |
| PoTableBackbone=gpt-4o-mini∼8B, Evaluation Protocol=Single inference run2025.01 | 88.93 | |
| H-Star2025.09 | 88.4 | |
| Chain-of-TableModel Category=Training-free, Pass@1=true2026.03 | 87.78 | |
| Qwen2.5-72B-InstructModel Category=Open-source, Pass@1=true2026.03 | 87.34 | |
| Table-R1Model Category=Tuning-based, Pass@1=true2026.03 | 87.17 | |
| PoTableBackbone=Llama-3.170B, Evaluation Protocol=Single inference run2025.01 | 87.06 | |
| TabLaP2025.09 | 86.9 | |
| Deepseek-R1Model Category=Open-source, Pass@1=true2026.03 | 86.25 | |
| DeBERTaV32022.11 | 86.2 | |
| Chain-of-TableBackbone=Llama-3.170B, Evaluation Protocol=Single inference run2025.01 | 85.62 | |
| SaMoE2022.11 | 85.1 | |
| REASTAP2022.10 | 84.7 | |
| BinderBackbone=gpt-4o-mini∼8B, Evaluation Protocol=Single inference run2025.01 | 84.63 | |
| Chain-of-TableBackbone=gpt-4o-mini∼8B, Evaluation Protocol=Single inference run2025.01 | 84.24 | |
| TAPEXpre-training_steps=50000, fine-tuning_steps=200002021.07 | 84.2 | |
| Tapex2022.11 | 84.2 | |
| Chain-of-Table2025.09 | 84.2 | |
| RePanda (Fact-Checking)Training=fine-tuned on PanTabFact2025.03 | 84.09 | |
| TAPEX2022.10 | 84 | |
| MM1.5-30BModel Scale=30B2024.09 | 84 | |
| GPT-4.1Model Category=Proprietary, Pass@1=true2026.03 | 83.94 | |
| T5Model Size=3B2022.10 | 83.7 | |
| TableMasterBackbone=gpt-3.5-turbo∼175B, Evaluation Protocol=Single inference run2025.01 | 83.65 | |
| DecompTaPas2022.10 | 82.7 | |
| TableLlama-7BModel Category=Tuning-based, Pass@1=true2026.03 | 82.53 | |
| MiniCPM-V-2.6 8BReasoning Strategy=HIPPO, # TIR=55.2K2026.02 | 82.27 | |
| Salience-aware learning systemMethod Type=Non-logical program-driven, Configuration=Full system2021.09 | 82.1 | |
| Salience-aware learning systemMethod Type=Non-logical program-driven, Configuration=w/o augmented data2021.09 | 82.1 | |
| Tree-of-TableBackbone=gpt-3.5-turbo∼175B, Evaluation Protocol=Single inference run2025.01 | 81.92 | |
| Salience-aware learning systemMethod Type=Non-logical program-driven, Configuration=w/o auxiliary task2021.09 | 81.9 | |
| TableFormer2022.10 | 81.6 | |
| DaterBackbone=Llama-3.170B, Evaluation Protocol=Single inference run2025.01 | 81.57 | |
| MATEIntermediate pre-training (CS)=true2021.09 | 81.4 | |
| Gemini-2.0-flashModel Category=Proprietary, Pass@1=true2026.03 | 81.03 | |
| TAPASModel=Large, Pre-training Objective=Counterfactual + Synthetic2020.10 | 81 | |
| Eisenschlos et al.model_category=Pre-trained Language Models2021.07 | 81 | |
| TaPas2022.10 | 81 | |
| Tapas2022.11 | 81 | |
| TAPASMethod Type=Non-logical program-driven, Reference=Eisenschlos et al., 20202021.09 | 81 | |
| TAPASIntermediate pre-training (CS)=true2021.09 | 81 | |
| DaterBackbone=gpt-4o-mini∼8B, Evaluation Protocol=Single inference run2025.01 | 80.98 | |
| BART2021.07 | 80.8 | |
| BART2022.10 | 80.5 | |
| mPLUG-DocOwl 1.5backbone=mPLUG-Owl 2, fine-tuning dataset=4M2024.11 | 80.4 | |
| DocOwl-1.5-ChatModel Scale=7B2024.09 | 80.2 | |
| Chain-of-TableBackbone=gpt-3.5-turbo∼175B, Evaluation Protocol=Single inference run2025.01 | 80.2 | |
| Dater2025.09 | 80.1 | |
| TabSQLifyBackbone=gpt-3.5-turbo∼175B, Evaluation Protocol=Single inference run2025.01 | 79.5 | |
| BinderBackbone=gpt-3.5-turbo∼175B, Evaluation Protocol=Single inference run2025.01 | 79.17 | |
| Tab-CoTModel Category=Training-free, Pass@1=true2026.03 | 79.05 | |
| TableGPT2-7BModel Category=Tuning-based, Pass@1=true2026.03 | 78.87 | |
| TabSQLify2025.09 | 78.8 | |
| TabSQLifyBackbone=gpt-4o-mini∼8B, Evaluation Protocol=Single inference run2025.01 | 78.75 | |
| TAPASModel=Base, Pre-training Objective=Counterfactual + Synthetic2020.10 | 78.5 | |
| BinderBackbone=Llama-3.170B, Evaluation Protocol=Single inference run2025.01 | 78.16 | |
| DaterBackbone=gpt-3.5-turbo∼175B, Evaluation Protocol=Single inference run2025.01 | 78.01 | |
| TAPASModel=Base, Pre-training Objective=Synthetic2020.10 | 77.9 | |
| Schlichtkrull et al.Retrieval=Oracle2022.11 | 77.6 | |
| Chain-of-Thought2025.09 | 77.2 | |
| Binder2025.09 | 77.2 | |
| MATEIntermediate pre-training (CS)=false2021.09 | 77 | |
| TAPASIntermediate pre-training (CS)=false2021.09 | 76.3 | |
| TAPASMethod Type=Non-logical program-driven, Reference=Dong and Smith, 20212021.09 | 76 | |
| MM1.5-7BModel Scale=7B2024.09 | 75.9 | |
| Qwen3-VL-8B-DISCOReasoning Strategy=Table-GLS, # TIR=02026.02 | 75.41 | |
| TAPASModel=Base, Pre-training Objective=Counterfactual2020.10 | 75.2 | |
| TAPASModel=Base, Pre-training Objective=SQA2020.10 | 74.6 | |
| Yang et al.model_category=Pre-trained Language Models2021.07 | 74.4 | |
| ProgVGAT2022.11 | 74.4 | |
| ProgVGATMethod Type=Logical program-driven2021.09 | 74.4 | |
| ReAcTableBackbone=gpt-3.5-turbo∼175B, Evaluation Protocol=Single inference run2025.01 | 74.4 | |
| End-to-End QA2025.09 | 73.5 | |
| Qwen2-VL-7B-Table-R1Reasoning Strategy=GRPO, # TIR=20.6K2026.02 | 73.4 | |
| Zhang et al.model_category=Pre-trained Language Models2021.07 | 73.2 | |
| SAT2022.10 | 73.2 | |
| SAT2022.11 | 73.2 | |
| MM1.5-3B-MOEModel Scale=3B, Architecture=MoE2024.09 | 73.1 | |
| MM1.5-3BModel Scale=3B2024.09 | 72.9 | |
| Shi et al.model_category=Pre-trained Language Models2021.07 | 72.3 | |
| HeterTFV2022.10 | 72.3 | |
| HeterTFVMethod Type=Logical program-driven2021.09 | 72.3 | |
| LOGICALFACTCHECKER (program from Seq2Action)2020.10 | 71.7 | |
| Zhong et al.model_category=Pre-trained Language Models2021.07 | 71.7 | |
| LFC2022.10 | 71.7 | |
| LogicFactChecker2022.11 | 71.7 | |
| LogicalFactCheckerMethod Type=Logical program-driven2021.09 | 71.7 | |
| LOGICALFACTCHECKER (program from LPA)2020.10 | 71.6 | |
| Few-Shot QABackbone=gpt-3.5-turbo∼175B, Evaluation Protocol=Single inference run2025.01 | 71.54 | |
| MM1.5-1B-MOEModel Scale=1B, Architecture=MoE2024.09 | 71.4 | |
| TabSQLifyBackbone=Llama-3.170B, Evaluation Protocol=Single inference run2025.01 | 70.7 | |
| End-to-End QABackbone=gpt-3.5-turbo∼175B, Evaluation Protocol=Single inference run2025.01 | 70.45 |