Mathematical Reasoning on CMATH (test)
89.7AccuracyQwen2.5-Math-72B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-Math-72Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 89.7 | |
| DeepSeekMath-RLSize=7B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 88.8 | |
| T2MBackbone=LLaDA2.1-mini (16B MoE), Decoding Strategy=greedy decoding, Temperature=0, block length=32, LOWPROB parameters=τ=0.3, Cmax=1, ρmax=0.252026.04 | 88.25 | |
| DeepSeekMath-RLSize=7B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 87.6 | |
| Qwen2-Math-72Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 86.4 | |
| GPT-4Size=-, Reasoning Mode=Chain-of-Thought, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 86 | |
| DeepSeek-LLM-ChatSize=67B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 85.4 | |
| Qwen2-72Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 85.4 | |
| Qwen2.5-Math-7Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 85 | |
| DeepSeekMath-InstructSize=7B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 84.6 | |
| DeepSeekMath-InstructSize=7B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 84.3 | |
| Qwen2-Math-7Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 83.2 | |
| Qwen2.5-Math-1.5Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 83 | |
| T2TBackbone=LLaDA2.1-mini (16B MoE), Decoding Strategy=greedy decoding, Temperature=0, block length=322026.04 | 82.33 | |
| DeepSeek-LLM-ChatSize=67B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 80.3 | |
| Qwen2-Math-1.5Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 79.6 | |
| DeepSeek-Coder-V2-Lite-Baseshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 77.8 | |
| Qwen2-7Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 76.7 | |
| Llama-3.1-70Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 75.5 | |
| GPT-3.5Size=-, Reasoning Mode=Chain-of-Thought, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 73.8 | |
| DeepSeekMath-Base-7Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 71.7 | |
| MetaMathSize=70B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Majority Vote (32 candidates)2024.02 | 70.9 | |
| InternLM2-Math-Base-20Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 65.9 | |
| WizardMath-v1.0Size=70B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Majority Vote (32 candidates)2024.02 | 65.4 | |
| Qwen2-1.5Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 55.6 | |
| ToRASize=34B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 53.4 | |
| Llama-3.1-8Bshot=6-shot, prompting=few-shot chain-of-thought2024.09 | 51.5 |