Financial Reasoning on FinQA
77.6AccuracyGPT-5 mini-high
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| GPT-5 mini-highReasoning Capability=with Reasoning, Domain Focus=General2026.03 | 77.6 | 77.1 | — | — | — | — | |
| QianfanHuijin-70BModel Scale=70B+, Mode=Thinking2025.12 | 77.1 | 86.2 | — | — | — | — | |
| ODA-Fin-RL-8BDomain Focus=Financial2026.03 | 73.3 | 74.6 | — | — | — | — | |
| Qwen3-4B-ThinkingReasoning Capability=with Reasoning, Domain Focus=General2026.03 | 72.4 | 72.5 | — | — | — | — | |
| Qwen3-32BReasoning Capability=without Reasoning, Domain Focus=General2026.03 | 72.4 | 74.7 | — | — | — | — | |
| DeepSeek-V3.2Model Scale=70B+, Mode=Thinking2025.12 | 72.2 | 83.82 | — | — | — | — | |
| Qwen3-8BReasoning Capability=without Reasoning, Domain Focus=General2026.03 | 72.2 | 71.5 | — | — | — | — | |
| CPDFAModel=Average across models2026.04 | 72.1 | — | — | — | 4.8 | — | |
| Gemini 2.5 Flash-LiteReasoning Capability=with Reasoning, Domain Focus=General2026.03 | 72 | 67.7 | — | — | — | — | |
| CPFAModel=Average across models2026.04 | 72 | — | — | — | 4.5 | — | |
| CPFFModel=Average across models2026.04 | 72 | — | — | — | 6.1 | — | |
| CPAAModel=Average across models2026.04 | 71.9 | — | — | — | 3.9 | — | |
| CPDAAModel=Average across models2026.04 | 71.9 | — | — | — | 3.9 | — | |
| ASCModel=Average across models2026.04 | 70.4 | — | — | — | 6.9 | — | |
| ESCModel=Average across models2026.04 | 70.4 | — | — | — | 9.1 | — | |
| ODA-Fin-SFT-8BDomain Focus=Financial2026.03 | 69.8 | 72.1 | — | — | — | — | |
| QianfanHuijin-8BModel Scale=8B, Mode=Thinking2025.12 | 68.3 | 81.47 | — | — | — | — | |
| Fin-R1Domain Focus=Financial2026.03 | 67.7 | 61.4 | — | — | — | — | |
| GenICLModel=Qwen3-8B, Selection strategy=top-12026.02 | 67.3 | — | — | — | — | — | |
| Dianjin-R1-7BDomain Focus=Financial2026.03 | 67.2 | 70.3 | — | — | — | — | |
| DICLModel=Qwen3-8B, Selection strategy=top-12026.02 | 67.1 | — | — | — | — | — | |
| theory-guided context selection strategyModel=Qwen3-8B, Selection strategy=top-12026.02 | 66.7 | — | — | — | — | — | |
| Qwen3-8BModel Scale=8B, Mode=Thinking2025.12 | 65.6 | 75.88 | — | — | — | — | |
| DeepSeek-R1Model Scale=70B+, Mode=Thinking2025.12 | 65.5 | 82.56 | — | — | — | — | |
| Qwen2.5-7B-InstructReasoning Capability=without Reasoning, Domain Focus=General2026.03 | 64.6 | 68 | — | — | — | — | |
| BM25Model=Qwen3-8B, Selection strategy=top-12026.02 | 64.2 | — | — | — | — | — | |
| TopicKModel=Qwen3-8B, Selection strategy=top-12026.02 | 64.2 | — | — | — | — | — | |
| Qwen3-235BModel Scale=70B+, Mode=Thinking2025.12 | 63.5 | 82.61 | — | — | — | — | |
| ZeroModel=Qwen3-8B, Selection strategy=top-1, contextual_information=none2026.02 | 60.1 | — | — | — | — | — | |
| Llama-3.1-8B-InstructReasoning Capability=without Reasoning, Domain Focus=General2026.03 | 57.6 | 56.9 | — | — | — | — | |
| DianJin-R1-7BModel Scale=8B, Mode=Thinking2025.12 | 56.7 | 71.85 | — | — | — | — | |
| Qwen2.5-7BReasoning Capability=without Reasoning, Domain Focus=General2026.03 | 52.4 | 47.8 | — | — | — | — | |
| GenICLModel=Llama-3.1-8B, Selection strategy=top-12026.02 | 51.6 | — | — | — | — | — | |
| DICLModel=Llama-3.1-8B, Selection strategy=top-12026.02 | 51.3 | — | — | — | — | — | |
| theory-guided context selection strategyModel=Llama-3.1-8B, Selection strategy=top-12026.02 | 51.1 | — | — | — | — | — | |
| TopicKModel=Llama-3.1-8B, Selection strategy=top-12026.02 | 49.8 | — | — | — | — | — | |
| BM25Model=Llama-3.1-8B, Selection strategy=top-12026.02 | 49.2 | — | — | — | — | — | |
| ZeroModel=Llama-3.1-8B, Selection strategy=top-1, contextual_information=none2026.02 | 48.6 | — | — | — | — | — | |
| Teacher (CoT)Size=32B2026.05 | 48.5 | — | — | — | — | — | |
| CoCoDASize=8B2026.05 | 40.16 | — | — | — | — | — | |
| ReToolSize=8B2026.05 | 39.65 | — | — | — | — | — | |
| ToRLSize=8B2026.05 | 38.19 | — | — | — | — | — | |
| Student (CoT)Size=8B2026.05 | 37.31 | — | — | — | — | — | |
| ReActSize=8B2026.05 | 34.69 | — | — | — | — | — | |
| CoCoDASize=4B2026.05 | 29.06 | — | — | — | — | — | |
| ReToolSize=4B2026.05 | 28.95 | — | — | — | — | — | |
| TroVESize=8B2026.05 | 28.9 | — | — | — | — | — | |
| CREATORSize=8B2026.05 | 27.29 | — | — | — | — | — | |
| TroVESize=4B2026.05 | 26.88 | — | — | — | — | — | |
| ToRLSize=4B2026.05 | 26.17 | — | — | — | — | — | |
| CREATORSize=4B2026.05 | 25.94 | — | — | — | — | — | |
| ReActSize=4B2026.05 | 25.82 | — | — | — | — | — | |
| Student (CoT)Size=4B2026.05 | 24.93 | — | — | — | — | — | |
| Xuanyuan-6B-ChatDomain Focus=Financial2026.03 | 23.4 | 36.7 | — | — | — | — | |
| CoCoDASize=1.7B2026.05 | 21.37 | — | — | — | — | — | |
| ReToolSize=1.7B2026.05 | 20.09 | — | — | — | — | — | |
| ToRLSize=1.7B2026.05 | 18.97 | — | — | — | — | — | |
| TroVESize=1.7B2026.05 | 18.66 | — | — | — | — | — | |
| CREATORSize=1.7B2026.05 | 16.72 | — | — | — | — | — | |
| ReActSize=1.7B2026.05 | 16.47 | — | — | — | — | — | |
| Student (CoT)Size=1.7B2026.05 | 16.04 | — | — | — | — | — | |
| CoCoDASize=0.6B2026.05 | 11.24 | — | — | — | — | — | |
| ReToolSize=0.6B2026.05 | 10.55 | — | — | — | — | — | |
| ToRLSize=0.6B2026.05 | 8.55 | — | — | — | — | — | |
| TroVESize=0.6B2026.05 | 8.49 | — | — | — | — | — | |
| CREATORSize=0.6B2026.05 | 7.89 | — | — | — | — | — | |
| ReActSize=0.6B2026.05 | 7.65 | — | — | — | — | — | |
| Student (CoT)Size=0.6B2026.05 | 5.93 | — | — | — | — | — | |
| Plutus-8B-InstructDomain Focus=Financial2026.03 | 5.5 | 36.2 | — | — | — | — | |
| 0-shot CoTBackbone=Llama-3.1-8B, Evaluation Protocol=0-shot2026.05 | — | — | — | — | — | 58 | |
| BERT Base Internal Retriever + DPR-FAISS External Retriever + RoBERTa Large Encoder GeneratorInternal Retriever=BERT Base, External Retriever=DPR-FAISS, Encoder Generator=RoBERTa Large, Training Epochs=202025.12 | — | — | 58.49 | 60.96 | — | — | |
| BERT Base Internal Retriever + RoBERTa Large Encoder Generator (Baseline)Internal Retriever=BERT Base, Encoder Generator=RoBERTa Large, Training Epochs=202025.12 | — | — | 57.87 | 60.02 | — | — | |
| DarkForestModels=Qwen2.5-7B-Instruct, finance-Llama3-8B, Saul-7B-Instruct-v12026.05 | — | — | 11.33 | 15.67 | — | — | |
| DebateModels=Qwen2.5-7B-Instruct, finance-Llama3-8B, Saul-7B-Instruct-v12026.05 | — | — | 1.33 | 3.67 | — | — | |
| Fino1, SFTBackbone=Llama-3.1-8B, Method Scaffold=SFT2026.05 | — | — | — | — | — | 60.87 | |
| Graph-of-Agent (Max)Models=Qwen2.5-7B-Instruct, finance-Llama3-8B, Saul-7B-Instruct-v12026.05 | — | — | 7 | 11.33 | — | — | |
| Graph-of-Agent (Mean)Models=Qwen2.5-7B-Instruct, finance-Llama3-8B, Saul-7B-Instruct-v12026.05 | — | — | 8.67 | 16 | — | — | |
| MAGEBackbone=Llama-3.1-8B, Execution Scaffold=Sonnet, Number of seeds=52026.05 | — | — | — | — | — | 67.5 | |
| MAGEBackbone=Llama-3.1-8B, Execution Scaffold=Opus, Number of seeds=52026.05 | — | — | — | — | — | 69 | |
| Mixture-of-AgentModels=Qwen2.5-7B-Instruct, finance-Llama3-8B, Saul-7B-Instruct-v12026.05 | — | — | 4.67 | 8.33 | — | — | |
| ReConcileModels=Qwen2.5-7B-Instruct, finance-Llama3-8B, Saul-7B-Instruct-v12026.05 | — | — | 1.33 | 4.33 | — | — | |
| RefineModels=Qwen2.5-7B-Instruct, finance-Llama3-8B, Saul-7B-Instruct-v12026.05 | — | — | 3.33 | 8.67 | — | — | |
| SecBERT Internal Retriever + BERT Large Encoder GeneratorInternal Retriever=SecBERT, Encoder Generator=BERT Large, Training Epochs=202025.12 | — | — | 53.95 | 55.76 | — | — | |
| SecBERT Internal Retriever + DPR-FAISS External Retriever + BERT Large Encoder GeneratorInternal Retriever=SecBERT, External Retriever=DPR-FAISS, Encoder Generator=BERT Large, Training Epochs=202025.12 | — | — | 54.64 | 56.97 | — | — | |
| SecBERT Internal Retriever + DPR-FAISS External Retriever + FinBERT Encoder GeneratorInternal Retriever=SecBERT, External Retriever=DPR-FAISS, Encoder Generator=FinBERT, Training Epochs=202025.12 | — | — | 48.77 | 50.9 | — | — | |
| SecBERT Internal Retriever + DPR-FAISS External Retriever + RoBERTa Base Encoder GeneratorInternal Retriever=SecBERT, External Retriever=DPR-FAISS, Encoder Generator=RoBERTa Base, Training Epochs=202025.12 | — | — | 56.75 | 58.81 | — | — | |
| SecBERT Internal Retriever + DPR-FAISS External Retriever + RoBERTa Large Encoder GeneratorInternal Retriever=SecBERT, External Retriever=DPR-FAISS, Encoder Generator=RoBERTa Large, Training Epochs=202025.12 | — | — | 60.54 | 63.48 | — | — | |
| SecBERT Internal Retriever + FinBERT Encoder GeneratorInternal Retriever=SecBERT, Encoder Generator=FinBERT, Training Epochs=202025.12 | — | — | 49.53 | 52.08 | — | — | |
| SecBERT Internal Retriever + RoBERTa Base Encoder GeneratorInternal Retriever=SecBERT, Encoder Generator=RoBERTa Base, Training Epochs=202025.12 | — | — | 56.49 | 58.33 | — | — | |
| SecBERT Internal Retriever + RoBERTa Large Encoder GeneratorInternal Retriever=SecBERT, Encoder Generator=RoBERTa Large, Training Epochs=202025.12 | — | — | 59.7 | 62.11 | — | — | |
| Self-ConsistencyModels=Qwen2.5-7B-Instruct, finance-Llama3-8B, Saul-7B-Instruct-v12026.05 | — | — | 3.33 | 6.33 | — | — |