Financial Question Answering on FinQA (test)
76.05AccuracyEEDP
Evaluation Results
| Method | Links | |
|---|---|---|
| EEDPBackbone=GPT-42024.02 | 76.05 | |
| Program-of-Thoughts (PoT)Backbone=GPT-42024.02 | 75.26 | |
| Chain-of-Thought (CoT)Backbone=GPT-42024.02 | 72.38 | |
| Program-of-Thoughts (PoT)Backbone=GPT-3.5-Turbo2024.02 | 68.97 | |
| DirectBackbone=GPT-42024.02 | 65.12 | |
| EEDPBackbone=PaLM 2-540B2024.02 | 61.95 | |
| EEDPBackbone=GPT-3.5-Turbo2024.02 | 61.88 | |
| ReMemModel=Qwen3-8B2026.02 | 61.5 | |
| Ours (theory-guided context selection strategy)Model=Qwen3-8B2026.02 | 61.4 | |
| ExpRAGModel=Qwen3-8B2026.02 | 61.1 | |
| DCModel=Qwen3-8B2026.02 | 60.9 | |
| BM25Model=Qwen3-8B2026.02 | 60.4 | |
| ZeroModel=Qwen3-8B2026.02 | 60.1 | |
| Chain-of-Thought (CoT)Backbone=GPT-3.5-Turbo2024.02 | 59.18 | |
| Ours (theory-guided context selection strategy)Model=Llama-3.1-8B2026.02 | 50 | |
| ExpRAGModel=Llama-3.1-8B2026.02 | 49.7 | |
| ReMemModel=Llama-3.1-8B2026.02 | 49.6 | |
| DCModel=Llama-3.1-8B2026.02 | 49.5 | |
| BM25Model=Llama-3.1-8B2026.02 | 49 | |
| ZeroModel=Llama-3.1-8B2026.02 | 48.6 | |
| DecomposersBackbone=PaLM 2-540B2024.02 | 46.38 | |
| TableMind++Model Category=Tuning-based, Pass@1=true2026.03 | 45.48 | |
| DecomposersBackbone=GPT-42024.02 | 44.93 | |
| TableMindModel Category=Tuning-based, Pass@1=true2026.03 | 42.02 | |
| Table-R1Model Category=Tuning-based, Pass@1=true2026.03 | 41.27 | |
| DirectBackbone=GPT-3.5-Turbo2024.02 | 40.47 | |
| Deepseek-R1Model Category=Open-source, Pass@1=true2026.03 | 37.42 | |
| Qwen2.5-72B-InstructModel Category=Open-source, Pass@1=true2026.03 | 35.83 | |
| Chain-of-Thought (CoT)Backbone=MAmmoTH-13B2024.02 | 35.32 | |
| EEDPBackbone=MAmmoTH-13B2024.02 | 35.05 | |
| EEDPBackbone=Mistral-7B2024.02 | 34.86 | |
| Chain-of-Thought (CoT)Backbone=PaLM 2-540B2024.02 | 34.79 | |
| Chain-of-Thought (CoT)Backbone=Mistral-7B2024.02 | 34.23 | |
| Gemini-2.5-flashModel Category=Proprietary, Pass@1=true2026.03 | 33.45 | |
| DecomposersBackbone=GPT-3.5-Turbo2024.02 | 32.33 | |
| EEDPBackbone=Llama 2-13B2024.02 | 30.47 | |
| Program-of-Thoughts (PoT)Backbone=PaLM 2-540B2024.02 | 30.41 | |
| DirectBackbone=PaLM 2-540B2024.02 | 30.33 | |
| Chain-of-TableModel Category=Training-free, Pass@1=true2026.03 | 29.47 | |
| GPT-5Model Category=Proprietary, Pass@1=true2026.03 | 28.93 | |
| Qwen3-8BModel Category=Open-source, Pass@1=true2026.03 | 27.83 | |
| Tab-CoTModel Category=Training-free, Pass@1=true2026.03 | 27.81 | |
| DirectBackbone=Mistral-7B2024.02 | 26.11 | |
| Chain-of-Thought (CoT)Backbone=Llama 2-13B2024.02 | 25.34 | |
| PoTableModel Category=Training-free, Pass@1=true2026.03 | 24.93 | |
| DirectBackbone=MAmmoTH-13B2024.02 | 22.83 | |
| Gemini-2.0-flashModel Category=Proprietary, Pass@1=true2026.03 | 19.62 | |
| TableGPT2-7BModel Category=Tuning-based, Pass@1=true2026.03 | 17.69 | |
| DecomposersBackbone=MAmmoTH-13B2024.02 | 17.65 | |
| Program-of-Thoughts (PoT)Backbone=MAmmoTH-13B2024.02 | 15.86 | |
| TableLlama-7BModel Category=Tuning-based, Pass@1=true2026.03 | 15.73 | |
| Program-of-Thoughts (PoT)Backbone=Llama 2-13B2024.02 | 12.97 | |
| DecomposersBackbone=Mistral-7B2024.02 | 12.34 | |
| DecomposersBackbone=Llama 2-13B2024.02 | 11.91 | |
| Program-of-Thoughts (PoT)Backbone=Mistral-7B2024.02 | 10.56 | |
| GPT-4.1Model Category=Proprietary, Pass@1=true2026.03 | 6.45 | |
| DirectBackbone=Llama 2-13B2024.02 | 1.8 |