Mathematical Reasoning on HMMT November 2025 (held-out)
49.11Pass@1ExpRL-Outcome
Evaluation Results
| Method | Links | |
|---|---|---|
| ExpRL-OutcomePriming=Dense terminal reward to full rollouts, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 49.11 | |
| ExpRL-ProcessPriming=Dense rewards to partial rollouts and prefixes, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 48.13 | |
| Self-DistillationPriming=Distillation-based, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 46.08 | |
| GRPOPriming=Verifiable sparse reward RL, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 42.91 | |
| Qwen3-4B-InstructModel type=Original base model, Evaluation=Sampling 128 responses per problem2026.06 | 40.6 | |
| SFTPriming=Supervised fine-tuning, Downstream RL=Stage-II RL setup, Evaluation=Sampling 128 responses per problem2026.06 | 20.09 |