Mathematical Reasoning on AIME 25 (Mean@8 accuracy)
73.95Mean@8 AccuracyMTI
Evaluation Results
| Method | Links | |
|---|---|---|
| MTIModel Series=Qwen3-14B2026.02 | 73.95 | |
| CCDModel Series=Qwen3-14B2026.02 | 73.75 | |
| CCDModel Series=Qwen3-8B2026.02 | 71.67 | |
| MTIModel Series=Qwen3-8B2026.02 | 69.46 | |
| BaseModel Series=Qwen3-14B, Decoding Mode=Standard decoding2026.02 | 68.75 | |
| CCDModel Series=Qwen3-4B2026.02 | 67.92 | |
| BaseModel Series=Qwen3-8B, Decoding Mode=Standard decoding2026.02 | 67.91 | |
| MTIModel Series=Qwen3-4B2026.02 | 66.81 | |
| BaseModel Series=Qwen3-4B, Decoding Mode=Standard decoding2026.02 | 63.75 | |
| GRPO w/ Ground-TruthRL Category=Ground-Truth RL (Oracle), Algorithm=GRPO, Reward Method=Ground-Truth2026.03 | 46.67 | |
| GRPO w/ Structure RewardRL Category=Label-free RL, Algorithm=GRPO, Reward Method=Structure Reward2026.03 | 45.83 | |
| GRPO w/ Entropy Minimization (EMPO)RL Category=Label-free RL, Algorithm=GRPO, Reward Method=Entropy Minimization (EMPO)2026.03 | 44.58 | |
| PPO w/ Structure RewardRL Category=Label-free RL, Algorithm=PPO, Reward Method=Structure Reward2026.03 | 42.92 | |
| GRPO w/ Majority Voting (TTRL)RL Category=Label-free RL, Algorithm=GRPO, Reward Method=Majority Voting (TTRL)2026.03 | 42.91 | |
| PPO w/ Ground-TruthRL Category=Ground-Truth RL (Oracle), Algorithm=PPO, Reward Method=Ground-Truth2026.03 | 41.67 | |
| PACED Forward KLDistillation track=Qwen3-14B → Qwen3-8B, Family=forward KL family2026.03 | 35.6 | |
| AKLDistillation track=Qwen3-14B → Qwen3-8B, Family=forward KL family2026.03 | 34.1 | |
| Hard Filter Forward KLDistillation track=Qwen3-14B → Qwen3-8B, Family=forward KL family2026.03 | 33.9 | |
| Base2026.03 | 31.67 | |
| Forward KL (unweighted)Distillation track=Qwen3-14B → Qwen3-8B, Family=forward KL family2026.03 | 29.3 | |
| BaseDistillation track=Qwen3-14B → Qwen3-8B, Family=forward KL family2026.03 | 20.8 |