Table Mathematical Reasoning on TabMWP (test)
99.57AccuracyTableMind++
Evaluation Results
| Method | Links | |
|---|---|---|
| TableMind++Model Category=Tuning-based, Pass@1=true2026.03 | 99.57 | |
| TableMindModel Category=Tuning-based, Pass@1=true2026.03 | 99.27 | |
| Deepseek-R1Model Category=Open-source, Pass@1=true2026.03 | 98.03 | |
| Qwen2.5-72B-InstructModel Category=Open-source, Pass@1=true2026.03 | 97.45 | |
| Chain-of-TableModel Category=Training-free, Pass@1=true2026.03 | 96.17 | |
| GPT-5Model Category=Proprietary, Pass@1=true2026.03 | 96.12 | |
| Table-R1Model Category=Tuning-based, Pass@1=true2026.03 | 96.02 | |
| MOCA-AGENTBackbone=Qwen3.6-27B, Method Category=Tool-augmented / agent-based methods2026.06 | 96 | |
| Gemini-2.5-flashModel Category=Proprietary, Pass@1=true2026.03 | 95.89 | |
| PoTableModel Category=Training-free, Pass@1=true2026.03 | 95.76 | |
| CREATORBackbone=ChatGPT, Method Category=Tool-augmented / agent-based methods2026.06 | 94.7 | |
| ChameleonBackbone=ChatGPT, Method Category=Tool-augmented / agent-based methods2026.06 | 93.28 | |
| TaCoBackbone=TAPEX-large, Method Category=Fine-tuning approaches2026.06 | 92.91 | |
| PoT+DocBackbone=ChatGPT, Method Category=Tool-augmented / agent-based methods2026.06 | 92.69 | |
| Tab-CoTModel Category=Training-free, Pass@1=true2026.03 | 92.21 | |
| Gemini-2.0-flashModel Category=Proprietary, Pass@1=true2026.03 | 91.56 | |
| CoTBackbone=GPT-4, Method Category=Prompting / chain-of-thought with LLMs2026.06 | 90.81 | |
| GPT-4.1Model Category=Proprietary, Pass@1=true2026.03 | 90.34 | |
| CoS-PlanningBackbone=ChatGPT, Method Category=Tool-augmented / agent-based methods2026.06 | 90 | |
| PoTBackbone=ChatGPT, Method Category=Prompting / chain-of-thought with LLMs2026.06 | 89.49 | |
| CRITICBackbone=ChatGPT, Method Category=Tool-augmented / agent-based methods2026.06 | 89 | |
| RetICLBackbone=Codex, Method Category=Prompting / chain-of-thought with LLMs2026.06 | 88.51 | |
| CRAFTBackbone=ChatGPT, Method Category=Tool-augmented / agent-based methods2026.06 | 88.4 | |
| CRITICBackbone=GPT-3, Method Category=Tool-augmented / agent-based methods2026.06 | 87.6 | |
| TaCoBackbone=TAPEX-base, Method Category=Fine-tuning approaches2026.06 | 86.12 | |
| TableLlama-7BModel Category=Tuning-based, Pass@1=true2026.03 | 83.88 | |
| CoTBackbone=ChatGPT, Method Category=Prompting / chain-of-thought with LLMs2026.06 | 82.03 | |
| PoT-SCBackbone=Codex, Method Category=Prompting / chain-of-thought with LLMs2026.06 | 81.8 | |
| TableGPT2-7BModel Category=Tuning-based, Pass@1=true2026.03 | 80.43 | |
| CRITICBackbone=Llama-2-70B, Method Category=Tool-augmented / agent-based methods2026.06 | 75 | |
| ToRABackbone=Llama-2-70B, Method Category=Fine-tuning approaches2026.06 | 74 | |
| ToRA-CodeBackbone=CodeLlama-34B, Method Category=Fine-tuning approaches2026.06 | 70.5 | |
| Qwen3-8BModel Category=Open-source, Pass@1=true2026.03 | 60.86 | |
| ToRABackbone=Llama-2-13B, Method Category=Fine-tuning approaches2026.06 | 47.2 | |
| ToRABackbone=Llama-2-7B, Method Category=Fine-tuning approaches2026.06 | 42.4 |