Mathematical Reasoning on ConceptMath Overall
78.02Average AccuracyGPT-4
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4Model Size=Proprietary2024.02 | 78.02 | |
| Qwen-72BModel Size=72B2024.02 | 70.73 | |
| GPT-3.5Model Size=Proprietary2024.02 | 70.58 | |
| InternLM2-Math-20BModel Size=20B2024.02 | 65.28 | |
| DeepSeekMath-7BModel Size=7B2024.02 | 64.02 | |
| InternLM2-Math-7BModel Size=7B2024.02 | 61.43 | |
| Yi-34BModel Size=34B2024.02 | 57.31 | |
| InternLM2-20BModel Size=20B2024.02 | 57.03 | |
| Qwen-14BModel Size=14B2024.02 | 54.36 | |
| InternLM2-7BModel Size=7B2024.02 | 54.08 | |
| Baichuan2-13BModel Size=13B2024.02 | 53.65 | |
| ChatGLM3-6BModel Size=6B2024.02 | 49.9 | |
| Yi-6BModel Size=6B2024.02 | 49.49 | |
| MAmmoTH-13BModel Size=13B2024.02 | 41.42 | |
| LLaMA2-70BModel Size=70B2024.02 | 39.81 | |
| LLaMA2-13BModel Size=13B2024.02 | 36.37 | |
| MetaMath-13BModel Size=13B2024.02 | 35.21 | |
| LLaMA2-7BModel Size=7B2024.02 | 30.57 | |
| WizardMath-13BModel Size=13B2024.02 | 27.7 |