Mathematical Reasoning on AIME 2026 (held-out) (Pass@1)
63.41Pass@1ExpRL-Process
Evaluation Results
| Method | Links | |
|---|---|---|
| ExpRL-ProcessPriming=Dense rewards to partial rollouts and prefixes, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 63.41 | |
| ExpRL-OutcomePriming=Dense terminal reward to full rollouts, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 61.74 | |
| GRPOPriming=Verifiable sparse reward RL, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 58.75 | |
| Self-DistillationPriming=Distillation-based, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 58.41 | |
| Qwen3-4B-InstructModel type=Original base model, Evaluation=Sampling 128 responses per problem2026.06 | 51.4 | |
| SFTPriming=Supervised fine-tuning, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 30.26 |