Mathematical Reasoning on MATH (Topic Breakdown)
2.3Algebra AccuracyFairSeq
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| FairSeqModel Size=125M, Evaluation Protocol=5-shot2022.04 | 2.3 | 0.8 | 0 | 1 | 1.9 | 1.3 | 0.2 | |
| FairSeqModel Size=6.7B, Evaluation Protocol=5-shot2022.04 | 1.7 | 1.5 | 1.5 | 1.1 | 2.8 | 2.1 | 0.2 | |
| FairSeqModel Size=2.7B, Evaluation Protocol=5-shot2022.04 | 1.4 | 1.7 | 1.5 | 1 | 1.1 | 1.1 | 0 | |
| FairSeqModel Size=1.3B, Evaluation Protocol=5-shot2022.04 | 1.3 | 1.5 | 0.6 | 0.7 | 0.7 | 1 | 0.4 | |
| GPT-JParameters=6B, Zero-shot=true2022.04 | 1.3 | 1.1 | 0.4 | 0.4 | 0.7 | 1 | 0.5 | |
| FairSeqModel Size=13B, Evaluation Protocol=5-shot2022.04 | 1.2 | 1.7 | 0.6 | 0.4 | 1.9 | 1.3 | 0 | |
| FairSeqModel Size=355M, Evaluation Protocol=5-shot2022.04 | 1 | 0.4 | 1.3 | 0.2 | 0.9 | 0.8 | 0.2 | |
| GPT-NeoXParameters=20B, Zero-shot=true2022.04 | 1 | 1.7 | 1.7 | 0.1 | 1.3 | 1.8 | 0.5 | |
| GPT-3 BabbageZero-shot=true2022.04 | 0.8 | 0.4 | 0 | 0.3 | 0 | 0.6 | 0 | |
| GPT-3 DaVinciZero-shot=true2022.04 | 0.8 | 0.6 | 0.2 | 0.3 | 1.1 | 1.4 | 0.4 | |
| GPT-3 AdaZero-shot=true2022.04 | 0.3 | 0 | 0 | 0 | 0.7 | 0.7 | 0.4 | |
| GPT-3 CurieZero-shot=true2022.04 | 0.3 | 0 | 0.2 | 0.6 | 0.6 | 0.8 | 0.2 |