Mathematical Reasoning on GSM8k (Pass@1)
95.2Pass@1Qwen2.5-Math-7B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-Math-7B-InstructZero-shot=true2025.01 | 95.2 | |
| o1-miniModel Type=Proprietary, Evaluation Protocol=zero-shot2025.01 | 94.8 | |
| GPT-4oModel Type=Proprietary, Evaluation Protocol=zero-shot2025.01 | 90.5 | |
| CoR-Math-7BModel Type=GMM, Evaluation Protocol=zero-shot2025.01 | 88.7 | |
| CoR-Math-7BZero-shot=true2025.01 | 88.7 | |
| DeepSeekMath-RL-7BZero-shot=true2025.01 | 88.2 | |
| GPT-4Model Type=Proprietary, Evaluation Protocol=zero-shot2025.01 | 87.1 | |
| DART-Math-7BZero-shot=true2025.01 | 86.6 | |
| InternLM2-Math-Plus-7BModel Type=GMM, Evaluation Protocol=zero-shot2025.01 | 85.8 | |
| Xwin-Math-7BZero-shot=true2025.01 | 82.6 | |
| MetaMath-Mistral-7BZero-shot=true2025.01 | 77.7 | |
| Llama-3.1-8B-InstructModel Type=Foundation, Evaluation Protocol=zero-shot2025.01 | 76.6 | |
| Mixtral-8x7BModel Type=Foundation, Evaluation Protocol=few-shot2025.01 | 74.4 | |
| DeepSeekMath-Instruct-7BZero-shot=true2025.01 | 73.6 | |
| NuminaMath-7B-TIRZero-shot=true, Source=Reported results with open-sourced weights2025.01 | 73.6 | |
| ToRA-CODEZero-shot=true2025.01 | 72.6 | |
| MetaMath-Llemma-7BZero-shot=true2025.01 | 69.2 | |
| ToRAZero-shot=true2025.01 | 68.8 | |
| NuminaMath-7B-CoTZero-shot=true, Source=Reported results with open-sourced weights2025.01 | 66.6 | |
| MetaMath-7BZero-shot=true2025.01 | 66.5 | |
| WizardMathZero-shot=true2025.01 | 54.9 | |
| InternLM2-Math-7B-BaseModel Type=GMM, Evaluation Protocol=few-shot2025.01 | 49.2 | |
| Llemma-7BModel Type=GMM, Evaluation Protocol=few-shot2025.01 | 41 | |
| Mistral-7BModel Type=Foundation, Evaluation Protocol=few-shot2025.01 | 40.3 | |
| MUSTARDModel Type=GMM, Evaluation Protocol=zero-shot2025.01 | 27.9 | |
| DeepSeekMath-7B-BaseModel Type=GMM, Evaluation Protocol=zero-shot2025.01 | 22.2 | |
| Llama-3.1-8BModel Type=Foundation, Evaluation Protocol=zero-shot2025.01 | 6.2 |