Mathematical Reasoning on AIME 2025 (held-out) (Pass@1)
59.07Pass@1ExpRL-Outcome
Evaluation Results
| Method | Links | |
|---|---|---|
| ExpRL-OutcomePriming=Dense terminal reward to full rollouts, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 59.07 | |
| ExpRL-ProcessPriming=Dense rewards to partial rollouts and prefixes, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 58.08 | |
| GRPOPriming=Verifiable sparse reward RL, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 55.99 | |
| Self-DistillationPriming=Distillation-based, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 55.59 | |
| Qwen3-4B-InstructModel type=Original base model, Evaluation=Sampling 128 responses per problem2026.06 | 46.46 | |
| HSDBase Model=Qwen3-32B-Base2026.06 | 41.7 | |
| GRPO + OPSDBase Model=Qwen3-32B-Base2026.06 | 38.6 | |
| RLSDBase Model=Qwen3-32B-Base2026.06 | 38.2 | |
| SDPOBase Model=Qwen3-32B-Base2026.06 | 37.8 | |
| OPSDBase Model=Qwen3-32B-Base2026.06 | 37.1 | |
| DAPOBase Model=Qwen3-32B-Base2026.06 | 36.7 | |
| GSPOBase Model=Qwen3-32B-Base2026.06 | 36 | |
| Dr. GRPOBase Model=Qwen3-32B-Base2026.06 | 35.4 | |
| GRPOBase Model=Qwen3-32B-Base2026.06 | 34.5 | |
| HSDBase Model=Qwen3-8B-Base2026.06 | 28.1 | |
| SFTPriming=Supervised fine-tuning, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 26.62 | |
| GRPO + OPSDBase Model=Qwen3-8B-Base2026.06 | 25.7 | |
| RLSDBase Model=Qwen3-8B-Base2026.06 | 25.3 | |
| SDPOBase Model=Qwen3-8B-Base2026.06 | 25 | |
| OPSDBase Model=Qwen3-8B-Base2026.06 | 24.4 | |
| DAPOBase Model=Qwen3-8B-Base2026.06 | 23.7 | |
| GSPOBase Model=Qwen3-8B-Base2026.06 | 23.2 | |
| Dr. GRPOBase Model=Qwen3-8B-Base2026.06 | 22.5 | |
| GRPOBase Model=Qwen3-8B-Base2026.06 | 21.6 |