Table Question Answering on WTQ (test)
64.7Denotation AccuracyT5 + TC + P
Evaluation Results
| Method | Links | |
|---|---|---|
| T5 + TC + PModel Category=Text-to-SQL, Table Content=true, Post-processing=true2024.09 | 64.7 | |
| OmniTabModel Category=E2E TQA2024.09 | 62.6 | |
| Readi-GPT3.5category=Inference-based Method, LLM=gpt-3.5-turbo2024.03 | 61.7 | |
| Readi-GPT4category=Inference-based Method, LLM=GPT42024.03 | 61.3 | |
| DocLayLLMSize(B)=82026.03 | 58.6 | |
| DocCogitoSize(B)=82026.03 | 58.3 | |
| TAPEXcategory=Training-based Method2024.03 | 57.5 | |
| Qwen3-VL-InstructSize(B)=82026.03 | 57.1 | |
| GPT4category=Inference-based Method2024.03 | 57 | |
| TAPEXModel Category=E2E TQA2024.09 | 57 | |
| GPT3.5category=Inference-based Method2024.03 | 55.8 | |
| DocCogitoSize(B)=42026.03 | 54.8 | |
| FRESBase Model=Pixtral 12B2025.05 | 54.4 | |
| Qwen3-VL-InstructSize(B)=42026.03 | 54.4 | |
| MM1.5-30BModel Scale=30B2024.09 | 54.1 | |
| PixtralModel Size=12B, Representation=text and image2025.05 | 52.5 | |
| MartenSize(B)=8.12026.03 | 52.4 | |
| StructGPTcategory=Inference-based Method2024.03 | 52.2 | |
| PixtralModel Size=12B, Representation=text2025.05 | 51.3 | |
| UnifiedSKGcategory=Training-based Method, backbone=T5-3B2024.03 | 49.3 | |
| TAPAScategory=Training-based Method2024.03 | 48.8 | |
| DocFormerv2Size(B)=0.752026.03 | 48.3 | |
| TabPediaInput Size=25602024.06 | 47.8 | |
| Phi-3-Vision-4BModel Scale=4B2024.09 | 47.4 | |
| TextHawk2Size(B)=7.42026.03 | 46.2 | |
| MM1.5-7BModel Scale=7B2024.09 | 46 | |
| GPT4VInput Size=6452024.06 | 45.5 | |
| PixtralModel Size=12B, Representation=image2025.05 | 42.2 | |
| MM1.5-3BModel Scale=3B2024.09 | 41.8 | |
| DocOwl-1.5-ChatModel Scale=7B2024.09 | 40.6 | |
| mPLUG-DocOwl 1.5backbone=mPLUG-Owl 2, fine-tuning dataset=4M2024.11 | 40.6 | |
| DocOwl-1.5-ChatSize(B)=8.12026.03 | 40.6 | |
| DocOwl 1.5Input Size=13442024.06 | 39.8 | |
| DocOwl-1.5Size(B)=8.12026.03 | 39.8 | |
| MM1.5-3B-MOEModel Scale=3B, Architecture=MoE2024.09 | 39.1 | |
| MM1.5-1B-MOEModel Scale=1B, Architecture=MoE2024.09 | 38.9 | |
| FRESBase Model=TableLlaVA 7B2025.05 | 38.8 | |
| Zhang et al.Size(B)=8.12026.03 | 38.6 | |
| TextMonkeyInput Size=8962024.06 | 37.9 | |
| DocOwl2Size(B)=82026.03 | 36.5 | |
| InternVL2-2BModel Scale=1B2024.09 | 35.8 | |
| Davinci-003category=Inference-based Method2024.03 | 34.8 | |
| HVFAbackbone=mPLUG-Owl, fine-tuning dataset=650K2024.11 | 34.5 | |
| Park et al.Size(B)=7.22026.03 | 34.5 | |
| TableLlaVAModel Size=7B, Representation=text2025.05 | 34.4 | |
| MM1.5-1BModel Scale=1B2024.09 | 34.1 | |
| MM1-30BModel Scale=30B2024.09 | 33.3 | |
| DocKylinSize(B)=7.12026.03 | 32.4 | |
| Gemini ProInput Size=6592024.06 | 32.3 | |
| HRVDASize(B)=7.12026.03 | 31.2 | |
| CogagentInput Size=11202024.06 | 30.2 | |
| MM1-7BModel Scale=7B2024.09 | 28.8 | |
| Xcomposer2Input Size=5112024.06 | 28.7 | |
| MonkeyInput Size=8962024.06 | 25.3 | |
| MonkeySize(B)=9.82026.03 | 25.3 | |
| MiniCPM-V 2.0-3BModel Scale=3B2024.09 | 24.2 | |
| MM1-3BModel Scale=3B2024.09 | 23.6 | |
| MM1-1BModel Scale=1B2024.09 | 19.9 | |
| TableLlaVAModel Size=7B, Representation=image2025.05 | 19.2 | |
| DonutSize(B)=0.262026.03 | 18.8 | |
| TableLlaVAModel Size=7B, Representation=text and image2025.05 | 17.2 | |
| SPHINX-TinyModel Scale=1B2024.09 | 15.3 |