Mathematical Reasoning on IMO-AnswerBench (held-out)
37.85Pass@1 ScoreExpRL-Outcome
Evaluation Results
| Method | Links | |
|---|---|---|
| ExpRL-OutcomePriming=Dense terminal reward to full rollouts, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 37.85 | |
| ExpRL-ProcessPriming=Dense rewards to partial rollouts and prefixes, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 35.73 | |
| GRPOPriming=Verifiable sparse reward RL, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 35.28 | |
| Self-DistillationPriming=Distillation-based, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 35.18 | |
| Qwen3-4B-InstructModel type=Original base model, Evaluation=Sampling 128 responses per problem2026.06 | 31.37 | |
| SFTPriming=Supervised fine-tuning, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 21.8 |