Fact Verification on TabFact
92.1AccuracyHuman
Evaluation Results
| Method | Links | |
|---|---|---|
| Human2024.04 | 92.1 | |
| TabTrim-8BMethod Category=Ours, Backbone/Model=TabTrim-8B, Pipeline Reasoner=GPT-4o-mini2026.01 | 91.2 | |
| PASTA2024.04 | 90.8 | |
| Table-CriticMethod Category=Critique methods, Pipeline Reasoner=GPT-4o-mini2026.01 | 90.6 | |
| TabTrim-4BMethod Category=Ours, Backbone/Model=TabTrim-4B, Pipeline Reasoner=GPT-4o-mini2026.01 | 89.4 | |
| Chain-of-TableMethod Category=LLM-based methods, Pipeline Reasoner=GPT-4o-mini2026.01 | 88.9 | |
| TALONMethod Category=Critique methods, Pipeline Reasoner=GPT-4o-mini2026.01 | 87.6 | |
| SaMoE2024.04 | 86.7 | |
| CHAIN-OF-TABLELLM Backbone=PaLM 22024.01 | 86.61 | |
| TAPEX2024.04 | 85.9 | |
| DATERLLM=Codex2024.04 | 85.6 | |
| IXC 2.5Size=7B, Visual tokens=51182024.09 | 85.2 | |
| BINDERLLM=Codex2024.04 | 85.1 | |
| DaterLLM Backbone=PaLM 22024.01 | 84.63 | |
| GRABBackbone=Qwen3-4B-Base2026.06 | 84.25 | |
| Liu et al. (2021)Extra Pre-training=true2022.01 | 84.2 | |
| TaPas2024.04 | 83.9 | |
| T5-3BBackbone=T5-3B2022.01 | 83.68 | |
| DaterMethod Category=LLM-based methods, Pipeline Reasoner=GPT-4o-mini2026.01 | 83.6 | |
| BinderMethod Category=Program-based methods, Pipeline Reasoner=GPT-4o-mini2026.01 | 83.3 | |
| ReAcTableLLM=Codex2024.04 | 83.1 | |
| TableLlama+textLLM=Llama-2 7B, Representation=OCR textual table2026.02 | 82.55 | |
| Soft PromptBackbone=Qwen3-4B-Base2026.06 | 81.15 | |
| T5-largeBackbone=T5-large2022.01 | 80.85 | |
| CHAIN-OF-TABLELLM Backbone=GPT 3.52024.01 | 80.2 | |
| Chain-of-Table2024.04 | 80.2 | |
| DocOwl 1.5Size=8B, Visual tokens=16982024.09 | 80.2 | |
| TabSQLifyMode=col+row2024.04 | 79.5 | |
| BinderLLM Backbone=GPT 3.52024.01 | 79.17 | |
| BINDERLLM=chatgpt2024.04 | 79.1 | |
| DiVA-FormerModality=V+T2026.03 | 79.1 | |
| Chain-of-ThoughtLLM Backbone=PaLM 22024.01 | 79.05 | |
| TabSQLifyMode=row2024.04 | 78.5 | |
| TabSQLifyMethod Category=Program-based methods, Pipeline Reasoner=GPT-4o-mini2026.01 | 78.3 | |
| DocOwl2Size=8B, Visual tokens=3242024.09 | 78.2 | |
| Few-Shot QALLM Backbone=PaLM 22024.01 | 78.06 | |
| DaterLLM Backbone=GPT 3.52024.01 | 78.01 | |
| DATERLLM=chatgpt2024.04 | 78 | |
| End-to-End QALLM Backbone=PaLM 22024.01 | 77.92 | |
| GPT-4o-miniMethod Category=Direct QA, Backbone/Model=GPT-4o-mini2026.01 | 77.4 | |
| TabSQLifyMode=col2024.04 | 77 | |
| BinderLLM Backbone=PaLM 22024.01 | 76.98 | |
| Qwen3-8BMethod Category=Direct QA, Backbone/Model=Qwen3-8B2026.01 | 76.7 | |
| T5-baseBackbone=T5-base2022.01 | 76.13 | |
| SAT2024.04 | 75.5 | |
| Yang et al. (2020)Extra Pre-training=false2022.01 | 74.4 | |
| LogicFactChecker2024.04 | 74.3 | |
| Qwen3-4BMethod Category=Direct QA, Backbone/Model=Qwen3-4B2026.01 | 74.1 | |
| ReAcTableLLM=chatgpt2024.04 | 73.1 | |
| TableCoTLLM=chatgpt2024.04 | 73.1 | |
| DirectModality=V+T2026.03 | 72.8 | |
| TableCoTLLM=Codex2024.04 | 72.6 | |
| DirectModality=Text2026.03 | 72.4 | |
| AdapterModality=V+T2026.03 | 72.4 | |
| DirectModality=Vision2026.03 | 72 | |
| Few-Shot QALLM Backbone=GPT 3.52024.01 | 71.54 | |
| AdapterModality=Vision2026.03 | 70.8 | |
| End-to-End QALLM Backbone=GPT 3.52024.01 | 70.45 | |
| Text-to-SQLLLM Backbone=PaLM 22024.01 | 68.37 | |
| Table-BERT2024.04 | 68.1 | |
| AdapterModality=Text2026.03 | 68 | |
| BaseBackbone=Qwen3-4B-Base2026.06 | 67.86 | |
| UReaderSize=7B, Visual tokens=8412024.09 | 67.6 | |
| CHAIN-OF-TABLELLM Backbone=LLaMA 22024.01 | 67.24 | |
| Chain-of-ThoughtLLM Backbone=GPT 3.52024.01 | 65.37 | |
| DaterLLM Backbone=LLaMA 22024.01 | 65.12 | |
| Text-to-SQLLLM Backbone=GPT 3.52024.01 | 64.71 | |
| Text-to-SQLLLM Backbone=LLaMA 22024.01 | 64.03 | |
| BinderLLM Backbone=LLaMA 22024.01 | 62.76 | |
| Few-Shot QALLM Backbone=LLaMA 22024.01 | 62.01 | |
| ResamplerModality=Vision2026.03 | 60.8 | |
| Chain-of-ThoughtLLM Backbone=LLaMA 22024.01 | 60.52 | |
| DocOwlSize=7B, Visual tokens=8412024.09 | 60.2 | |
| Table-LLaVA 7BLLM=Vicuna-1.5 7B2026.02 | 59.85 | |
| Re-Table-7B-rerankLLM=Mistral 7B, Method Variant=rerank2026.02 | 56.67 | |
| ResamplerModality=V+T2026.03 | 56.4 | |
| ResamplerModality=Text2026.03 | 51.6 | |
| Re-Table-7B-retrievalLLM=Mistral 7B, Method Variant=retrieval2026.02 | 50.64 | |
| End-to-End QALLM Backbone=LLaMA 22024.01 | 44.86 | |
| MonkeyLLM=Qwen 7B2026.02 | 22.56 | |
| LLaVA v1.6LLM=Vicuna-1.5 7B2026.02 | 19.26 | |
| LLaVA v1.5LLM=Vicuna-1.5 7B2026.02 | 18.9 | |
| BLIP2LLM=Flan-T5 7B2026.02 | 18.62 | |
| Vary-toyLLM=Qwen 1.8B2026.02 | 6.33 | |
| Llama2+textLLM=Llama-2 7B, Representation=OCR textual table2026.02 | 4.21 | |
| MiniGPT-4LLM=Vicuna 7B2026.02 | 0 |