Code Generation on MathQA Python Original (test)
84.7Pass@80GPT-Neo 125M
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-Neo 125Mmethod=self-sampling FCS + PCS2022.05 | 84.7 | 77.6 | |
| LaMDA 137Bpretrained on code=false2022.05 | 81.2 | — | |
| LaMDA 68Bpretrained on code=false2022.05 | 79.5 | — | |
| Codex Davincifew-shot learning=true2022.05 | 42 | 6 |