Arithmetic Reasoning on ASDiv
93.74AccuracyOriginal
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| OriginalBackbone=Qwen3-8b2025.10 | 93.74 | — | — | |
| UCoT-0.9Backbone=Qwen3-8b, Compression Ratio=0.92025.10 | 93.53 | — | — | |
| Automatic Model Selection with LLMsBackbone=GPT-4, Decoding strategy=Greedy2023.05 | 93.5 | — | — | |
| LightThinkerBackbone=Qwen3-8b2025.10 | 93.42 | — | — | |
| UCoT-0.7Backbone=Qwen3-8b, Compression Ratio=0.72025.10 | 93.32 | — | — | |
| UCoT-0.5Backbone=Qwen3-8b, Compression Ratio=0.52025.10 | 92.77 | — | — | |
| CoTBackbone=GPT-4, Decoding strategy=Greedy2023.05 | 92.7 | — | — | |
| PALBackbone=GPT-4, Decoding strategy=Greedy2023.05 | 90.2 | — | — | |
| Automatic Model Selection with LLMsBackbone=ChatGPT, Decoding strategy=Greedy2023.05 | 89.4 | — | — | |
| CoTBackbone=ChatGPT, Decoding strategy=Greedy2023.05 | 89.3 | — | — | |
| DIVERSEModel=code-davinci-0022022.06 | 88.7 | — | — | |
| CoconutBackbone=Qwen3-8b2025.10 | 88.265 | — | — | |
| Self-consistencyBackbone=GPT-3 Code-davinci-0022022.03 | 87.8 | — | — | |
| Depth-RecurrentBackbone=Qwen3-8b2025.10 | 87.316 | — | — | |
| OracleDescription=paired upper bound2026.04 | 87 | — | — | |
| Self-ConsistencyModel=code-davinci-0022022.06 | 86.2 | — | — | |
| TAGPolicy=fit-selected2026.04 | 85.2 | 7.7 | 3.8 | |
| DIVERSEModel=text-davinci-0022022.06 | 83.5 | — | — | |
| PALBackbone=ChatGPT, Decoding strategy=Greedy2023.05 | 83 | — | — | |
| Self-ConsistencyModel=PaLM 540B2022.06 | 81.9 | — | — | |
| Self-consistencyBackbone=PaLM-540B2022.03 | 81.9 | — | — | |
| Automatic Model Selection with LLMsBackbone=Codex, Decoding strategy=Greedy2023.05 | 81.6 | — | — | |
| CodexEngine=code-davinci-002, Prompting=Chain of thought, External Calculator=false2022.01 | 80.4 | — | — | |
| CoTBackbone=Codex, Decoding strategy=Greedy2023.05 | 80.2 | — | — | |
| CoT-promptingBackbone=GPT-3 Code-davinci-0022022.03 | 80.1 | — | — | |
| CodexEngine=code-davinci-002, Prompting=Chain of thought, External Calculator=true2022.01 | 80 | — | — | |
| PALBackbone=Codex, Decoding strategy=Greedy2023.05 | 79.1 | — | — | |
| BaselineDescription=single-pass decoding2026.04 | 77.5 | — | — | |
| RetryDescription=compute-matched second pass2026.04 | 77.5 | — | — | |
| Self-ConsistencyModel=text-davinci-0022022.06 | 76.9 | — | — | |
| Greedy DecodeModel=code-davinci-0022022.06 | 75.5 | — | — | |
| Prior bestPrompting=N/A2022.01 | 75.3 | — | — | |
| SOTA (Fine-tuned)Mode=Fine-tuned2022.06 | 75.3 | — | — | |
| Lan et al. (2021)2022.03 | 75.3 | — | — | |
| CodexEngine=code-davinci-002, Prompting=Standard, External Calculator=false2022.01 | 74 | — | — | |
| Greedy DecodeModel=PaLM 540B2022.06 | 74 | — | — | |
| CoT-promptingBackbone=PaLM-540B2022.03 | 74 | — | — | |
| PaLM 540BPrompting=Chain of thought, External Calculator=false2022.01 | 73.9 | — | — | |
| PaLM 540BPrompting=Chain of thought, External Calculator=true2022.01 | 72.6 | — | — | |
| PaLM 540BPrompting=Standard, External Calculator=false2022.01 | 72.1 | — | — | |
| GPT-3 175BEngine=text-davinci-002, Prompting=Chain of thought, External Calculator=false2022.01 | 71.3 | — | — | |
| GPT-3 175BEngine=text-davinci-002, Prompting=Chain of thought, External Calculator=true2022.01 | 71.1 | — | — | |
| GPT-3 175BEngine=text-davinci-002, Prompting=Standard, External Calculator=false2022.01 | 70.3 | — | — | |
| Self-consistencyBackbone=GPT-3 Code-davinci-0012022.03 | 61.9 | — | — | |
| Greedy DecodeModel=text-davinci-0022022.06 | 60.8 | — | — | |
| Self-ConsistencyModel=LaMDA 137B2022.06 | 58.2 | — | — | |
| Self-consistencyBackbone=LaMDA-137B2022.03 | 58.2 | — | — | |
| DIVERSEModel=GPT-3 davinci (175B)2022.06 | 57.6 | — | — | |
| LaMDA 137BPrompting=Chain of thought, External Calculator=true2022.01 | 53.4 | — | — | |
| Self-ConsistencyModel=GPT-3 davinci (175B)2022.06 | 52.8 | — | — | |
| CoT-promptingBackbone=GPT-3 Code-davinci-0012022.03 | 52.7 | — | — | |
| Greedy DecodeModel=LaMDA 137B2022.06 | 49 | — | — | |
| CoT-promptingBackbone=LaMDA-137B2022.03 | 49 | — | — | |
| LaMDA 137BPrompting=Chain of thought, External Calculator=false2022.01 | 46.6 | — | — | |
| UL2 20BPrompting=CoT prompting, Calculator=True, Self-consistency=True2022.05 | 43.5 | — | — | |
| LaMDA 137BPrompting=Standard, External Calculator=false2022.01 | 40.1 | — | — | |
| UL2 20BPrompting=Chain of thought, External Calculator=true2022.01 | 34.3 | — | — | |
| UL2 20BPrompting=CoT prompting, Calculator=True2022.05 | 34.3 | — | — | |
| Greedy DecodeModel=GPT-3 davinci (175B)2022.06 | 31.4 | — | — | |
| Self-consistencyBackbone=UL2-20B2022.03 | 21.5 | — | — | |
| ForgettingBase model=LLaMA-2-13B2025.08 | 17.8 | — | — | |
| UL2 20BPrompting=Chain of thought, External Calculator=false2022.01 | 16.9 | — | — | |
| UL2 20BPrompting=CoT prompting2022.05 | 16.9 | — | — | |
| CoT-promptingBackbone=UL2-20B2022.03 | 16.9 | — | — | |
| UL2 20BPrompting=Standard, External Calculator=false2022.01 | 16 | — | — | |
| UL2 20BPrompting=standard prompting2022.05 | 16 | — | — | |
| IgnoringBase model=LLaMA-2-13B2025.08 | 15.34 | — | — | |
| Full Tokens (standard SFT)Base model=LLaMA-2-13B2025.08 | 8.76 | — | — | |
| BaseBase model=LLaMA-2-13B2025.08 | 0.35 | — | — |