Mathematical Reasoning on CMath (pass@1)
91.7Pass@1Qwen2.5-7B + GRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-7B + GRPOBase Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=false2025.09 | 91.7 | |
| Qwen2.5-7B + GRPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=GRPO, VERL Integration=true2025.09 | 90.7 | |
| Qwen2.5-7B + PPO w/ VERL.Base Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=true2025.09 | 90.3 | |
| VCRDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 90.2 | |
| Qwen2.5-7B + PPOBase Model=Qwen2.5-7B, RL Framework=PPO, VERL Integration=false2025.09 | 89.8 | |
| MTTeacher Model=Qwen2.5-Math-7B-Instruct2026.04 | 89.8 | |
| DistillLM-2Teacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 89.7 | |
| DistilLLMTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 89.3 | |
| GKDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 88.8 | |
| ABKDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 88.5 | |
| KDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 88.2 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B, RL Framework=None, VERL Integration=false2025.09 | 86.7 | |
| GPT-42023.10 | 86 | |
| KwaiYiiMath-HPA#Params=13B, Human preference alignment=true2023.10 | 85.83 | |
| KwaiYiiMath#Params=13B, Human preference alignment=false2023.10 | 85.33 | |
| Ernie Bot2023.10 | 84.33 | |
| ChatGPT2023.10 | 73.83 | |
| ChatGLM2#Params=6B2023.10 | 68.36 | |
| Qwen2.5-DenseLLM=Qwen2.5, Model=Dense, Activated Parameters=1.5 B2026.05 | 63.99 | |
| QWen#Params=7B2023.10 | 63.16 | |
| MSStudent Model=Qwen2.5-Math-1.5B2026.04 | 62 | |
| Qwen2.5-Dense2MoELLM=Qwen2.5, Model=Ours, Activated Parameters=1.2 B2026.05 | 55.5 | |
| BaiChuan1#Params=13B2023.10 | 51.33 | |
| WizardMath#Params=13B2023.10 | 50.83 | |
| Llama-3.2-3B-Instruct + GRPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=true2025.09 | 46.2 | |
| Llama-3.2-3B-Instruct + PPO w/ VERL.Base Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=true2025.09 | 30.7 | |
| Llama-3.2-3B-Instruct + GRPOBase Model=Llama-3.2-3B-Instruct, RL Framework=GRPO, VERL Integration=false2025.09 | 28.3 | |
| Llama2-Dense2MoELLM=Llama2, Model=Ours, Activated Parameters=5.7 B2026.05 | 25.33 | |
| Llama-3.2-3B-Instruct + PPOBase Model=Llama-3.2-3B-Instruct, RL Framework=PPO, VERL Integration=false2025.09 | 21.2 | |
| Llama2-DenseLLM=Llama2, Model=Dense, Activated Parameters=7.0 B2026.05 | 16.6 | |
| Llama-3.2-3B-InstructBase Model=Llama-3.2-3B-Instruct, RL Framework=None, VERL Integration=false2025.09 | 10.2 |