Math Word Problem Solution Generation on GSM8K (test)
84.5ACC-opLEVER
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LEVERBase model=codex-davinci-002, Verification strategy=LEVER, Verifier Base model=RoBERTa-large2023.02 | 84.5 | — | — | — | |
| DiVeRSeFinetuning status=With Finetuning, Base model=Codex, Decoding strategy=DiVeRSe, Note=Generating natural language solutions instead of programs2023.02 | 83.2 | — | — | — | |
| PoT-SCFinetuning status=Without Finetuning, Base model=Program-of-Thought, Decoding strategy=Self-Consistency (SC)2023.02 | 80 | — | — | — | |
| Codex + SCFinetuning status=Without Finetuning, Base model=Codex, Decoding strategy=Self-Consistency (SC)2023.02 | 78 | — | — | — | |
| EP + MLBase model=codex-davinci-002, Decoding/Verification strategy=EP + ML2023.02 | 72.6 | — | — | — | |
| PALFinetuning status=Without Finetuning, Base model=PAL2023.02 | 72 | — | — | — | |
| GreedyBase model=codex-davinci-002, Decoding strategy=Greedy2023.02 | 67.2 | — | — | — | |
| Planning-T5-largeParameters=770M, Operation classifier=Yes2023.06 | 66.3 | 40.5 | 62.3 | 21.2 | |
| Planning-GPT-2-mediumParameters=345M, Operation classifier=Yes2023.06 | 65.2 | 39.5 | 61.8 | 20.1 | |
| Chain-of-thought-tuning T5-largeParameters=770M2023.06 | 63.1 | 35.3 | 58.9 | 17 | |
| Planning-GPT-2Parameters=117M, Operation classifier=Yes2023.06 | 61.6 | 35.4 | 56.7 | 14.1 | |
| Chain-of-thought-tuning GPT-2-mediumParameters=345M2023.06 | 61.1 | 38.1 | 58.1 | 16.1 | |
| Planning-T5Parameters=220M, Operation classifier=Yes2023.06 | 60.6 | 34.4 | 55.7 | 13.9 | |
| Chain-of-thought-tuning GPT-2Parameters=117M2023.06 | 55.1 | 34.3 | 49.4 | 8.1 | |
| Chain-of-thought-tuning T5Parameters=220M2023.06 | 52.1 | 30.3 | 45.4 | 3.1 | |
| Neo-1.3B + SCFinetuning status=With Finetuning, Base model=GPT-Neo-1.3B, Decoding strategy=SC2023.02 | 24.2 | — | — | — | |
| Neo-2.7B + SSFinetuning status=With Finetuning, Base model=GPT-Neo-2.7B, Decoding strategy=SS2023.02 | 19.5 | — | — | — |