Mathematical Reasoning on CMATH
95.7AccuracyQwen2.5-Math-72B-Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 95.7 | — | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 95.3 | — | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 94.5 | — | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 94.3 | — | |
| Qwen2-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 94.2 | — | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=rm@82024.09 | 94.2 | — | |
| Qwen2-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 94 | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 94 | — | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=rm@82024.09 | 93.8 | — | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=maj@82024.09 | 93.5 | — | |
| Qwen2-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 93.2 | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=rm@82024.09 | 93.2 | — | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=Greedy2024.09 | 93 | — | |
| Qwen2-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 92.8 | — | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 92.7 | — | |
| GPT-4o-2024-08-06Evaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 92.5 | — | |
| Qwen2-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 92.2 | — | |
| Qwen2-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 92.2 | — | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=maj@82024.09 | 92 | — | |
| CESBackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 91.9 | 391 | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 91.8 | — | |
| Qwen2-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 91.7 | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 91.7 | — | |
| FP8Model=7B2025.08 | 91.7 | — | |
| FP16Model=7B2025.08 | 91.5 | — | |
| MicroMixModel=7B2025.08 | 91.5 | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=maj@82024.09 | 90.8 | — | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=Greedy2024.09 | 90.5 | — | |
| DAPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 90.3 | 292 | |
| Entropy ShapeBackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 90.1 | 291 | |
| Qwen2-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 90 | — | |
| DeepSeek-Coder-V2-Lite-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 89.8 | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 89.7 | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=Greedy2024.09 | 89.3 | — | |
| R1-7BBackbone=DeepSeek-R1-Distill-Qwen-7B2026.05 | 89.2 | 292 | |
| Qwen2-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 88 | — | |
| NuminaMath-72B-CoTEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 87.3 | — | |
| DeepSeekMath-7B-RLEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 86.7 | — | |
| Llama-3.1-70B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 86.7 | — | |
| Internlm2-math-plus-mixtral8x7BEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 85.7 | — | |
| CESBackbone=DeepSeek-R1-Distill-1.5B2026.05 | 85.3 | 676 | |
| DAPOBackbone=DeepSeek-R1-Distill-1.5B2026.05 | 84.7 | 616 | |
| Qwen2-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 84.2 | — | |
| Qwen2-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 83.5 | — | |
| HSA-ULTraining Strategy=Annealing, Architecture=MoE, Total Params=8B, Activated Params=1B, Training Tokens=8T2025.11 | 82.88 | — | |
| Internlm2-math-plus-7BEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 82.7 | — | |
| Internlm2-math-plus-20BEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 81.3 | — | |
| NuminaMath-7B-CoTEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 78.2 | — | |
| AGPOclipping=adaptive, ATS=true2026.05 | 77.6 | — | |
| Mathstral-7B-v0.1Evaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 76.7 | — | |
| AGPOclipping=adaptive, ATS=false2026.05 | 76.4 | — | |
| GRPOATS=true2026.05 | 75.9 | — | |
| TRM-MoETraining Strategy=Base, Architecture=MoE, Total Params=8B, Activated Params=1B, Training Tokens=8T2025.11 | 74.59 | — | |
| HSA-ULTraining Strategy=Base, Architecture=MoE, Total Params=8B, Activated Params=1B, Training Tokens=8T2025.11 | 74.13 | — | |
| GRPOclipping=fixed ε2026.05 | 74.1 | — | |
| Adaptive-KL PPO2026.05 | 73.4 | — | |
| DPO2026.05 | 72.5 | — | |
| PPO2026.05 | 71.2 | — | |
| Qwen3Training Strategy=Annealing, Architecture=Dense, Total Params=0.6B, Activated Params=0.6B, Training Tokens=36T2025.11 | 66.67 | — | |
| Qwen2-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 65.5 | — | |
| Llama-3.1-8B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 64.8 | — | |
| HSA-ULTraining Strategy=Annealing, Architecture=Dense, Total Params=0.5B, Activated Params=0.5B, Training Tokens=4T2025.11 | 60.75 | — | |
| Qwen2.5Training Strategy=Annealing, Architecture=Dense, Total Params=0.5B, Activated Params=0.5B, Training Tokens=18T2025.11 | 52.09 | — |