Table Question Answering on WTQ
91.25AccuracyGemini-3-Pro-Preview
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini-3-Pro-PreviewReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 91.25 | — | — | |
| Claude-Opus-4.5Reasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 89.91 | — | — | |
| Doubao-1.5-thinking-proReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 88.25 | — | — | |
| GPT-4oReasoning Mode=No reasoning, Model Category=No reasoning Models2026.01 | 85.64 | — | — | |
| Qwen3-32B-InstructReasoning Mode=thinking, Model Category=Reasoning Models2026.01 | 85.62 | — | — | |
| QWQ-32BReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 84.99 | — | — | |
| Deepseek-R1Reasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 83.92 | — | — | |
| Qwen3-8B-Think-SFT-TabCodeRLReasoning Mode=thinking, Model Category=Table-Specific Models2026.01 | 83.07 | — | — | |
| OpenAI o1-miniReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 80.57 | — | — | |
| Qwen2.5-72B-InstructReasoning Mode=No reasoning, Model Category=No reasoning Models2026.01 | 77.45 | — | — | |
| Qwen3-8B-InstructReasoning Mode=thinking, Model Category=Reasoning Models2026.01 | 76.94 | — | — | |
| TabLaP-EWLearning Paradigm=Proposed (Ours)2024.10 | 76.6 | 77.6 | — | |
| TabLaPLearning Paradigm=Proposed (Ours)2024.10 | 76.6 | 77.9 | — | |
| Qwen2-72B-InstructScale=72B, Modality=Text2025.02 | 72.55 | — | 57.11 | |
| Mix-SCLearning Paradigm=Zero-shot LLM2024.10 | 72.5 | — | — | |
| Chain-of-TableLearning Paradigm=Few-shot LLM2024.10 | 67.3 | — | — | |
| Qwen3-8B-InstructReasoning Mode=no-thinking, Model Category=No reasoning Models2026.01 | 66.03 | — | — | |
| DATERLearning Paradigm=Few-shot LLM2024.10 | 65.9 | — | — | |
| GPT-5.4-miniSetting=Others2026.06 | 65.88 | — | — | |
| BinderLearning Paradigm=Few-shot LLM2024.10 | 64.6 | — | — | |
| OmniTab-LargeLearning Paradigm=Fine-tuned PLM2024.10 | 63.3 | — | — | |
| Qwen2.5-VL-Instruct + HIPPOScale=7B, Modality=Image & Text2025.02 | 62.73 | — | 62.9 | |
| GPT-4oScale=N/A, Modality=Image & Text2025.02 | 62.5 | — | 55.58 | |
| TableGPT2-7BReasoning Mode=Table-Specific, Model Category=Table-Specific Models2026.01 | 62.15 | — | — | |
| Qwen2.5-VL-Instruct + Vanilla SFTScale=7B, Modality=Image & Text2025.02 | 61.44 | — | 62.15 | |
| Qwen2.5-VL-InstructScale=7B, Modality=Image & Text2025.02 | 60.45 | — | 58.58 | |
| GRABSetting=Tuned LLM (LoRA)2026.06 | 58.75 | — | — | |
| GRABSetting=Frozen LLM2026.06 | 58.33 | — | — | |
| GPT-4oLearning Paradigm=Zero-shot LLM2024.10 | 58.1 | — | — | |
| Qwen2.5-InstructScale=7B, Modality=Text2025.02 | 57.91 | — | 61.54 | |
| zero-shotSetting=Tuned LLM (LoRA)2026.06 | 57.51 | — | — | |
| TAPEX-LargeLearning Paradigm=Fine-tuned PLM2024.10 | 57.5 | — | — | |
| TAMOSetting=Frozen LLM2026.06 | 56.95 | — | — | |
| Qwen2.5-VL-Instruct (Image Only)Scale=7B, Modality=Image2025.02 | 56.85 | — | 56.85 | |
| GPT-4o (Image Only)Scale=N/A, Modality=Image2025.02 | 56.39 | — | 54.26 | |
| Prompt tuningSetting=Frozen LLM2026.06 | 56.02 | — | — | |
| Llama-3.1-8B-InstructReasoning Mode=No reasoning, Model Category=No reasoning Models2026.01 | 55.9 | — | — | |
| MiniCPM-V-2.6 + HIPPOScale=8B, Modality=Image & Text2025.02 | 55.77 | — | 65.46 | |
| MiniCPM-V-2.6 + Vanilla SFTScale=8B, Modality=Image & Text2025.02 | 55.54 | — | 61.93 | |
| IXC 2.5Size=7B, Visual tokens=51182024.09 | 53.6 | — | — | |
| MiniCPM-V-2.6Scale=8B, Modality=Image & Text2025.02 | 52.3 | — | 62.12 | |
| GPT-3.5 TurboLearning Paradigm=Zero-shot LLM2024.10 | 50.9 | — | — | |
| TableLLamaSetting=Others2026.06 | 48.59 | — | — | |
| MiniCPM-V-2.6 (Image Only)Scale=8B, Modality=Image2025.02 | 47.97 | — | 60.42 | |
| Deepseek-R1-Distill-Qwen-7BReasoning Mode=Reasoning Models, Model Category=Reasoning Models2026.01 | 47.39 | — | — | |
| UDOPPipeline=specialized2023.05 | 47.2 | — | — | |
| Table-LLaVA + HIPPOScale=7B, Modality=Image & Text2025.02 | 41.39 | — | 55.52 | |
| DocOwl 1.5Size=8B, Visual tokens=16982024.09 | 40.6 | — | — | |
| Table-LLaVAScale=7B, Modality=Image & Text2025.02 | 38.28 | — | 50.98 | |
| TableGPT2IFT Data=2.36M, Backbone=Qwen2.5-7B2025.06 | 38.26 | — | — | |
| BARTMode=Text Baseline2023.05 | 38 | — | — | |
| DocOwl2Size=8B, Visual tokens=3242024.09 | 36.5 | — | — | |
| TableLlama + OracleScale=7B, Modality=Text2025.02 | 35.01 | — | 55.33 | |
| zero-shotSetting=Inference-only2026.06 | 32.29 | — | — | |
| TextMonkeySize=9B, Visual tokens=7682024.09 | 31.9 | — | — | |
| TableLLM-7BReasoning Mode=Table-Specific, Model Category=Table-Specific Models2026.01 | 31.74 | — | — | |
| Dublinvariable_resResolution=variable2023.05 | 29.7 | — | — | |
| UReaderSize=7B, Visual tokens=8412024.09 | 29.4 | — | — | |
| InternLM-XComposer2Model Generation Capability=Text only2024.07 | 28.7 | — | — | |
| TextHarmony-ChatModel Generation Capability=Text and image, Slide-LoRA=true2024.07 | 28.3 | — | — | |
| TextHarmonyModel Generation Capability=Text and image, Slide-LoRA=true2024.07 | 27.1 | — | — | |
| DocOwlSize=7B, Visual tokens=8412024.09 | 26.9 | — | — | |
| Llama3.1-InstructScale=8B, Modality=Text2025.02 | 26.31 | — | 35.42 | |
| TextHarmony*Model Generation Capability=Text and image, Slide-LoRA=false2024.07 | 26.1 | — | — | |
| Dublinfixed_resResolution=fixed2023.05 | 25.7 | — | — | |
| TableLlama-7BReasoning Mode=Table-Specific, Model Category=Table-Specific Models2026.01 | 25.65 | — | — | |
| MonkeySize=9B, Visual tokens=12802024.09 | 25.3 | — | — | |
| MonkeyModel Generation Capability=Text only2024.07 | 25.3 | — | — | |
| TableLlama + MarkdownScale=7B, Modality=Text2025.02 | 24.97 | — | 37.86 | |
| TableDreamerIFT Data=27K, Backbone=GPT-4o2025.06 | 22.88 | — | — | |
| TabPediaScale=7B, Modality=Image2025.02 | 21.71 | — | 20.84 | |
| GenQAIFT Data=100K2025.06 | 21.63 | — | — | |
| Llama3-InstructScale=8B, Modality=Text2025.02 | 21.24 | — | 31.98 | |
| MiniCPM3-4B2025.06 | 20.93 | — | — | |
| Table-LLaVA (Image Only)Scale=13B, Modality=Image2025.02 | 20.41 | — | 38.09 | |
| DynasourIFT Data=132K2025.06 | 20.11 | — | — | |
| MonkeyScale=7B, Modality=Image2025.02 | 19.07 | — | 14.16 | |
| OmniTabIFT Data=100K2025.06 | 18.84 | — | — | |
| DonutSize=<1B, Visual tokens=4800, Fine-tuned=true2024.09 | 18.8 | — | — | |
| Table-LLaVA (Image Only)Scale=7B, Modality=Image2025.02 | 18.43 | — | 35.69 | |
| TableDreamerIFT Data=27K, Backbone=Llama3.1-70B-Instruct2025.06 | 17.25 | — | — | |
| Mistral-7B-Instruct-v0.32025.06 | 16.41 | — | — | |
| Llama2Scale=7B, Modality=Text2025.02 | 16.39 | — | 17.53 | |
| TableLLMIFT Data=80K, Backbone=CodeLlama-7B2025.06 | 15.67 | — | — | |
| MM-InterleavedModel Generation Capability=Text and image2024.07 | 15.1 | — | — | |
| Yi-1.5-9B-Chat2025.06 | 14.02 | — | — | |
| TableLLM-syn-dataIFT Data=80K2025.06 | 13.92 | — | — | |
| MagpieIFT Data=100K2025.06 | 13.89 | — | — | |
| Self-InstructIFT Data=100K2025.06 | 13.77 | — | — | |
| InternLM2.5-7B-Chat2025.06 | 13.51 | — | — | |
| SEED-LLaMA-14BModel Generation Capability=Text and image2024.07 | 13.2 | — | — | |
| Evol-InstructIFT Data=100K2025.06 | 12.37 | — | — | |
| TableBenchLLMIFT Data=20K, Backbone=Llama3.1-8B2025.06 | 12.31 | — | — | |
| Llama3.1-8B-Instruct2025.06 | 11.35 | — | — | |
| ReasTapIFT Data=100K2025.06 | 9.96 | — | — | |
| TableGPT-syn-dataIFT Data=66K2025.06 | 9.13 | — | — | |
| GLM4-9B-Chat2025.06 | 8.94 | — | — | |
| UCTRIFT Data=43K2025.06 | 8.84 | — | — | |
| Phi3.5-mini-3.8B2025.06 | 7.99 | — | — | |
| Vary-toyScale=1.8B, Modality=Image2025.02 | 7.96 | — | 5.77 |