Arithmetic Reasoning on GSM8K (test)
97.35AccuracySGE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SGEModel=GPT-4, Tool=Code Interpreter2024.05 | 97.35 | — | |
| Automatic Model Selection with LLMsBackbone=GPT-4, Self-Consistency Paths=152023.05 | 96.8 | — | |
| Automatic Model Selection with LLMsBackbone=GPT-4, Self-Consistency Paths=52023.05 | 96.5 | — | |
| CoTBackbone=GPT-4, Self-Consistency Paths=152023.05 | 95.8 | — | |
| CoTBackbone=GPT-4, Self-Consistency Paths=52023.05 | 95.6 | — | |
| PALBackbone=GPT-4, Self-Consistency Paths=152023.05 | 95.5 | — | |
| PALBackbone=GPT-4, Self-Consistency Paths=52023.05 | 94.7 | — | |
| Self-AgreementModel=GPT-4, Scenario=Third scenario (type and format known)2023.11 | 93.9 | — | |
| Self-ConsistencyModel=GPT-4, Scenario=Third scenario (type and format known)2023.11 | 93.3 | — | |
| Decomp PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 91.85 | — | |
| Few-Shot CoTModel=GPT-4, Scenario=Third scenario (type and format known)2023.11 | 90.8 | — | |
| Refine PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 89.8 | — | |
| CoT PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 89.76 | — | |
| Automatic Model Selection with LLMsBackbone=ChatGPT, Self-Consistency Paths=152023.05 | 89.2 | — | |
| Automatic Model Selection with LLMsBackbone=ChatGPT, Self-Consistency Paths=52023.05 | 88.2 | — | |
| CoT + VotingLLM=GPT-3.5-turbo, Prompting Format=Chain-of-Thought (CoT), Voting Strategy=Majority Voting, k=102023.06 | 87.62 | — | |
| CoTBackbone=ChatGPT, Self-Consistency Paths=152023.05 | 87.4 | — | |
| Ours (Natural Program (NP), No Verification)LLM=GPT-3.5-turbo, Prompting Format=Natural Program (NP), Verification Method=None, Voting Strategy=Majority Voting, k=102023.06 | 87.05 | — | |
| IO PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 87.04 | — | |
| Ours (NP + Deductive Verification + UPV)LLM=GPT-3.5-turbo, Prompting Format=Natural Program (NP), Verification Method=Deductive Verification, Voting Strategy=UPV, k=102023.06 | 86.01 | — | |
| CoTBackbone=ChatGPT, Self-Consistency Paths=52023.05 | 85.4 | — | |
| PALBackbone=ChatGPT, Self-Consistency Paths=152023.05 | 82.4 | — | |
| Self-AgreementModel=GPT-3.5-turbo, Scenario=Scenario 32023.11 | 82.4 | — | |
| PALBackbone=ChatGPT, Self-Consistency Paths=52023.05 | 80.9 | — | |
| GPT-3.5-turboPrompting Strategy=Self-Consistency, Scenario=Scenario 32023.11 | 80.3 | — | |
| LLMBOOSTBackbone=Qwen-2.5-7B, Params=3 × 7B2025.12 | 78.9 | — | |
| MinervaPrompting Strategy=Self-Consistency, Scenario=Scenario 32023.11 | 78.5 | — | |
| LLMBOOSTBackbone=Qwen-2.5-7B, Params=2 × 7B2025.12 | 78.1 | — | |
| UNITEBackbone=Qwen-2.5-7B, Params=3 × 7B2025.12 | 77.8 | — | |
| voteBackbone=Qwen-2.5-7B, Params=3 × 7B2025.12 | 77.4 | — | |
| GPT-3.5-turboPrompting Strategy=USC, Scenario=Scenario 32023.11 | 76.8 | — | |
| voteBackbone=Qwen-2.5-7B, Params=2 × 7B2025.12 | 76.3 | — | |
| UNITEBackbone=Qwen-2.5-7B, Params=2 × 7B2025.12 | 76.3 | — | |
| Faithful CoT + VotingLLM=GPT-3.5-turbo, Prompting Format=Faithful CoT, Voting Strategy=Majority Voting, k=102023.06 | 75.8 | — | |
| T-copilotBackbone=Qwen-2.5-7B, Params=2 × 7B2025.12 | 75.6 | — | |
| Qwen-2.5-7BBackbone=Qwen-2.5-7B, Params=1 × 7B2025.12 | 75.1 | — | |
| PaLMPrompting Strategy=Self-Consistency, Scenario=Scenario 32023.11 | 74.4 | — | |
| LLMBOOSTBackbone=Qwen-2.5-3B, Params=2 × 3B2025.12 | 74 | — | |
| LLMBOOSTBackbone=Qwen-2.5-3B, Params=3 × 3B2025.12 | 73.4 | — | |
| T-copilotBackbone=Qwen-2.5-3B, Params=2 × 3B2025.12 | 72.1 | — | |
| voteBackbone=Qwen-2.5-3B, Params=3 × 3B2025.12 | 71.9 | — | |
| UNITEBackbone=Qwen-2.5-3B, Params=2 × 3B2025.12 | 71.6 | — | |
| UNITEBackbone=Qwen-2.5-3B, Params=3 × 3B2025.12 | 71.6 | — | |
| voteBackbone=Qwen-2.5-3B, Params=2 × 3B2025.12 | 71.2 | — | |
| GPT-3.5-turboPrompting Strategy=Few-Shot CoT, Scenario=Scenario 32023.11 | 70 | — | |
| Qwen-2.5-3BBackbone=Qwen-2.5-3B, Params=1 × 3B2025.12 | 69.5 | — | |
| LLMBOOSTBackbone=Llama-3.1-8B, Params=2 × 8B2025.12 | 68.8 | — | |
| LLMBOOSTBackbone=Llama-3.1-8B, Params=3 × 8B2025.12 | 68.5 | — | |
| voteBackbone=Llama-3.1-8B, Params=2 × 8B2025.12 | 67.8 | — | |
| T-copilotBackbone=Llama-3.1-8B, Params=2 × 8B2025.12 | 66.5 | — | |
| voteBackbone=Llama-3.1-8B, Params=3 × 8B2025.12 | 66 | — | |
| UNITEBackbone=Llama-3.1-8B, Params=3 × 8B2025.12 | 65.6 | — | |
| UNITEBackbone=Llama-3.1-8B, Params=2 × 8B2025.12 | 65.5 | — | |
| Llama-3.1-8BBackbone=Llama-3.1-8B, Params=1 × 8B2025.12 | 63.7 | — | |
| DoRA + MoDEKBackbone=LLaMA3-8B2024.10 | 63.1 | — | |
| DoRABackbone=LLaMA3-8B2024.10 | 62.7 | — | |
| LoRAALL + MoDEKBackbone=LLaMA3-8B2024.10 | 62.2 | — | |
| LoRAALLBackbone=LLaMA3-8B2024.10 | 61.6 | — | |
| LoRA¬K + MoDEK (+0.04%)Backbone=LLaMA3-8B2024.10 | 61.5 | — | |
| Self-AgreementModel=Llama-2-70B-Chat, Scenario=Third scenario (type and format known)2023.11 | 61 | — | |
| LoRA¬KBackbone=LLaMA3-8B2024.10 | 60.6 | — | |
| Fine-tunedBits=16, Model=LLaMA-3.1-8B2025.09 | 59.74 | — | |
| Self-ConsistencyModel=Llama-2-70B-Chat, Scenario=Third scenario (type and format known)2023.11 | 59.7 | — | |
| MinervaPrompting Strategy=Few-Shot CoT, Scenario=Scenario 32023.11 | 58.8 | — | |
| PaLM 540B (CoT 8-shot)Model=PaLM 540B, Finetuning Strategy=None (Zero-shot/Few-shot), Prompting Strategy=8-shot Chain-of-thought2022.12 | 56.9 | 58.6 | |
| PaLMPrompting Strategy=Few-Shot CoT, Scenario=Scenario 32023.11 | 56.5 | — | |
| QWHABits=4, Model=LLaMA-3.1-8B, Adapter Type=WHA, QA Init.=✓, Coefficient Selection=AdaAlloc2025.09 | 56.1 | — | |
| QWHABits=3, Model=LLaMA-3.1-8B, Adapter Type=WHA, QA Init.=✓, Coefficient Selection=AdaAlloc2025.09 | 55.34 | — | |
| GPT-3 175B finetuned + verifierModel=GPT-3 175B, Scenario=Scenario 32023.11 | 55 | — | |
| Fine-tunedBits=16, Model=Mistral-7B-v0.32025.09 | 54.51 | — | |
| SHiRABits=4, Model=LLaMA-3.1-8B, Adapter Type=Sparse, QA Init.=✗, Coefficient Selection=Random2025.09 | 54.36 | — | |
| LoCABits=4, Model=LLaMA-3.1-8B, Adapter Type=DCA, QA Init.=✗, Coefficient Selection=LoCA2025.09 | 54.36 | — | |
| LLMBOOSTBackbone=LLama-3.2-3B, Params=3 × 3B2025.12 | 54.3 | — | |
| REFTBackbone=LLaMA3-8B2024.10 | 54 | — | |
| SSHBits=4, Model=LLaMA-3.1-8B, Adapter Type=DHA, QA Init.=✗, Coefficient Selection=SSH2025.09 | 53.98 | — | |
| CLoQBits=4, Model=LLaMA-3.1-8B, Adapter Type=LoRA, QA Init.=✓2025.09 | 53.83 | — | |
| CLoQBits=3, Model=LLaMA-3.1-8B, Adapter Type=LoRA, QA Init.=✓2025.09 | 53.75 | — | |
| QWHABits=4, Model=Mistral-7B-v0.3, Adapter Type=WHA, QA Init.=✓, Coefficient Selection=AdaAlloc2025.09 | 53.68 | — | |
| voteBackbone=LLama-3.2-3B, Params=3 × 3B2025.12 | 53.5 | — | |
| LoCABits=3, Model=LLaMA-3.1-8B, Adapter Type=DCA, QA Init.=✗, Coefficient Selection=LoCA2025.09 | 53.15 | — | |
| LoQTBit=4, Model scale=LLaMA-2-13B2024.05 | 52.9 | — | |
| UNITEBackbone=LLama-3.2-3B, Params=3 × 3B2025.12 | 52.9 | — | |
| ApiQBit=4, Model scale=LLaMA-2-13B2024.05 | 52.4 | — | |
| CLoQBits=4, Model=Mistral-7B-v0.3, Adapter Type=LoRA, QA Init.=✓2025.09 | 52.01 | — | |
| QLoRABit=4, Model scale=LLaMA-2-13B2024.05 | 51.6 | — | |
| LLMBOOSTBackbone=LLama-3.2-3B, Params=2 × 3B2025.12 | 51.5 | — | |
| LoRABit=16, Model scale=LLaMA-2-13B2024.05 | 51.3 | — | |
| LoftQBit=4, Model scale=LLaMA-2-13B2024.05 | 51.3 | — | |
| SHiRABits=4, Model=Mistral-7B-v0.3, Adapter Type=Sparse, QA Init.=✗, Coefficient Selection=Random2025.09 | 51.02 | — | |
| voteBackbone=LLama-3.2-3B, Params=2 × 3B2025.12 | 50.9 | — | |
| UNITEBackbone=LLama-3.2-3B, Params=2 × 3B2025.12 | 50.9 | — | |
| SSHBits=3, Model=LLaMA-3.1-8B, Adapter Type=DHA, QA Init.=✗, Coefficient Selection=SSH2025.09 | 50.34 | — | |
| LLama-3.2-3BBackbone=LLama-3.2-3B, Params=1 × 3B2025.12 | 49.7 | — | |
| T-copilotBackbone=LLama-3.2-3B, Params=2 × 3B2025.12 | 49.5 | — | |
| Few-Shot CoTModel=Llama-2-70B-Chat, Scenario=Third scenario (type and format known)2023.11 | 48.1 | — | |
| LoCABits=4, Model=Mistral-7B-v0.3, Adapter Type=DCA, QA Init.=✗, Coefficient Selection=LoCA2025.09 | 47.99 | — | |
| SSHBits=4, Model=Mistral-7B-v0.3, Adapter Type=DHA, QA Init.=✗, Coefficient Selection=SSH2025.09 | 47.99 | — | |
| QWHABits=3, Model=Mistral-7B-v0.3, Adapter Type=WHA, QA Init.=✓, Coefficient Selection=AdaAlloc2025.09 | 47.84 | — | |
| LORAModel=13B, #Params.=0.67%2024.08 | 47.5 | — | |
| SSHBits=3, Model=Mistral-7B-v0.3, Adapter Type=DHA, QA Init.=✗, Coefficient Selection=SSH2025.09 | 47.15 | — |