Mathematical Problem Solving on MATH (Sub-domain Breakdown)
0.362Overall AccuracyDeepSeekMath-Base
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| DeepSeekMath-BaseSize=7B, Evaluation=Chain-of-thought prompting, Source=Open-Source Base Model2024.02 | 0.362 | — | — | — | — | — | — | — | |
| MinervaSize=540B, Evaluation=Chain-of-thought prompting, Source=Closed-Source Base Model2024.02 | 0.336 | — | — | — | — | — | — | — | |
| MinervaSize=62B, Evaluation=Chain-of-thought prompting, Source=Closed-Source Base Model2024.02 | 0.276 | — | — | — | — | — | — | — | |
| InternLM2-20Bshots=4-shot, model_size_group=~20B2024.03 | 0.255 | — | — | — | — | — | — | — | |
| LlemmaSize=34B, Evaluation=Chain-of-thought prompting, Source=Open-Source Base Model2024.02 | 0.253 | — | — | — | — | — | — | — | |
| Qwen-14Bshots=4-shot, model_size_group=~20B2024.03 | 0.251 | — | — | — | — | — | — | — | |
| Mixtral-8x7B-v0.1shots=4-shot, model_size_group=~20B2024.03 | 0.227 | — | — | — | — | — | — | — | |
| InternLM2-7Bshots=4-shot, model_size_group=~7B2024.03 | 0.202 | — | — | — | — | — | — | — | |
| ChatGLM3-6B-Baseshots=4-shot, model_size_group=~7B2024.03 | 0.192 | — | — | — | — | — | — | — | |
| LlemmaSize=7B, Evaluation=Chain-of-thought prompting, Source=Open-Source Base Model2024.02 | 0.181 | — | — | — | — | — | — | — | |
| MistralSize=7B, Evaluation=Chain-of-thought prompting, Source=Open-Source Base Model2024.02 | 0.143 | — | — | — | — | — | — | — | |
| MinervaSize=7B, Evaluation=Chain-of-thought prompting, Source=Closed-Source Base Model2024.02 | 0.141 | — | — | — | — | — | — | — | |
| InternLM2-20B-Baseshots=4-shot, model_size_group=~20B2024.03 | 0.136 | — | — | — | — | — | — | — | |
| Qwen-7Bshots=4-shot, model_size_group=~7B2024.03 | 0.134 | — | — | — | — | — | — | — | |
| Mistral-7B-v0.1shots=4-shot, model_size_group=~7B2024.03 | 0.113 | — | — | — | — | — | — | — | |
| Baichuan2-13B-Baseshots=4-shot, model_size_group=~20B2024.03 | 0.101 | — | — | — | — | — | — | — | |
| InternLM2-7B-Baseshots=4-shot, model_size_group=~7B2024.03 | 0.089 | — | — | — | — | — | — | — | |
| Baichuan2-7B-Baseshots=4-shot, model_size_group=~7B2024.03 | 0.055 | — | — | — | — | — | — | — | |
| Llama2-13Bshots=4-shot, model_size_group=~20B2024.03 | 0.049 | — | — | — | — | — | — | — | |
| Llama2-7Bshots=4-shot, model_size_group=~7B2024.03 | 0.033 | — | — | — | — | — | — | — | |
| GPT-JParameters=6B, Few-shot settings=5-shot2022.04 | — | 0.032 | 0.036 | 0.027 | 0.024 | 0.044 | 0.052 | 0.013 | |
| GPT-NeoXParameters=20B, Few-shot settings=5-shot2022.04 | — | 0.049 | 0.03 | 0.015 | 0.021 | 0.065 | 0.057 | 0.027 |