Mathematical Reasoning on AIME 2025 (Acc@16, pass@16)
75Avg@16Dense Rollout Dense Training
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Dense Rollout Dense TrainingModel Backbone=Qwen3-14B, Generation Length Cutoff=37K, Training Data Size=53K Polaris examples2026.06 | 75 | — | 82.69 | |
| Sparsity setup (0.86) RolloutModel Backbone=Qwen3-14B, Sparsity Level=0.86, Generation Length Cutoff=37K, Training Data Size=53K Polaris examples2026.06 | 75 | — | 82.8 | |
| Sparsity setup (0.86) RolloutModel Backbone=Qwen3-4B, Sparsity Level=0.86, Generation Length Cutoff=37K, Training Data Size=53K Polaris examples2026.06 | 74.16 | — | 83.21 | |
| Sparsity setup (0.92) RolloutModel Backbone=Qwen3-8B, Sparsity Level=0.92, Generation Length Cutoff=37K, Training Data Size=33K Polaris examples2026.06 | 72.29 | — | 80.76 | |
| Dense Rollout Dense TrainingModel Backbone=Qwen3-4B, Generation Length Cutoff=37K, Training Data Size=53K Polaris examples2026.06 | 71.45 | — | 79.33 | |
| Sparsity setup (0.86) RolloutModel Backbone=Qwen3-8B, Sparsity Level=0.86, Generation Length Cutoff=37K, Training Data Size=33K Polaris examples2026.06 | 71.45 | — | 80.55 | |
| Dense Rollout Dense TrainingModel Backbone=Qwen3-8B, Generation Length Cutoff=37K, Training Data Size=33K Polaris examples2026.06 | 70.83 | — | 80 | |
| Sparsity setup (0.92) RolloutModel Backbone=Qwen3-4B, Sparsity Level=0.92, Generation Length Cutoff=37K, Training Data Size=53K Polaris examples2026.06 | 70 | — | 79.78 | |
| Sparsity setup (0.92) RolloutModel Backbone=Qwen3-1.7B, Sparsity Level=0.92, Generation Length Cutoff=37K, Training Data Size=53K Polaris examples2026.06 | 46.25 | — | 57.29 | |
| Sparse with LoRA Distillation (DistillSparse)Model Backbone=Qwen3-1.7B, Generation Length Cutoff=37K, Training Data Size=53K Polaris examples2026.06 | 43.54 | — | 55.62 | |
| Dense Rollout Dense TrainingModel Backbone=Qwen3-1.7B, Generation Length Cutoff=37K, Training Data Size=53K Polaris examples2026.06 | 43.22 | — | 54.4 | |
| Sparsity setup (0.86) RolloutModel Backbone=Qwen3-1.7B, Sparsity Level=0.86, Generation Length Cutoff=37K, Training Data Size=53K Polaris examples2026.06 | 41.45 | — | 54.36 | |
| TeacherBackbone=JustRL-1.5B2026.06 | 35.6 | — | — | |
| OPRD2026.06 | 34.6 | — | — | |
| OPD top-162026.06 | 34 | — | — | |
| OPD top-1Variant=sampled-token2026.06 | 33.5 | — | — | |
| StudentBackbone=R1-distill-1.5B2026.06 | 21.9 | — | — | |
| PowerFlowBackbone=Qwen2.5-Math-7B2026.03 | 14.4 | — | — | |
| GRPOBackbone=Qwen2.5-Math-7B2026.03 | 12.9 | — | — | |
| EMPOBackbone=Qwen2.5-Math-7B2026.03 | 12.3 | — | — | |
| InstructBackbone=Qwen2.5-Math-7B2026.03 | 12.3 | — | — | |
| TTRLBackbone=Qwen2.5-Math-7B2026.03 | 11.9 | — | — | |
| PowerFlowBackbone=Qwen2.5-Math-1.5B2026.03 | 10 | — | — | |
| PowerSamplingBackbone=Qwen2.5-Math-7B2026.03 | 10 | — | — | |
| Qwen2.5-32B-InstructBackbone=Qwen2.5-32B-Instruct2026.03 | 9.4 | — | — | |
| Low-tempBackbone=Qwen2.5-Math-7B2026.03 | 7.7 | — | — | |
| GRPOBackbone=Qwen2.5-Math-1.5B2026.03 | 6.7 | — | — | |
| One-shot EMBackbone=Qwen2.5-Math-7B2026.03 | 6.2 | — | — | |
| InstructBackbone=Qwen2.5-Math-1.5B2026.03 | 6 | — | — | |
| Format-onlyBackbone=Qwen2.5-Math-1.5B2026.03 | 5 | — | — | |
| EMPOBackbone=Qwen2.5-Math-1.5B2026.03 | 4.6 | — | — | |
| BaseBackbone=Qwen2.5-Math-7B2026.03 | 4.2 | — | — | |
| Low-tempBackbone=Qwen2.5-Math-1.5B2026.03 | 4 | — | — | |
| BaseBackbone=Qwen2.5-Math-1.5B2026.03 | 1.9 | — | — | |
| PowerFlowBackbone=Qwen2.5-1.5B2026.03 | 1.5 | — | — | |
| IntuitorBackbone=Qwen2.5-1.5B2026.03 | 0.8 | — | — | |
| Low-tempBackbone=Llama-3.2-3B-Instruct2026.03 | 0.6 | — | — | |
| Low-tempBackbone=Qwen2.5-1.5B2026.03 | 0.4 | — | — | |
| GRPOBackbone=Qwen2.5-1.5B2026.03 | 0.4 | — | — | |
| PowerFlowBackbone=Llama-3.2-3B-Instruct2026.03 | 0.4 | — | — | |
| IntuitorBackbone=Llama-3.2-3B-Instruct2026.03 | 0.2 | — | — | |
| BaseBackbone=Qwen2.5-1.5B2026.03 | 0 | — | — | |
| InstructBackbone=Qwen2.5-1.5B2026.03 | 0 | — | — | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.03 | 0 | — | — | |
| GRPOBackbone=Llama-3.2-3B-Instruct2026.03 | 0 | — | — | |
| Det. Trunc.Base Model=Qwen2.5-Math-7B2026.02 | — | 2.7 | 7.1 | |
| Det. Trunc.Base Model=Qwen3-8B2026.02 | — | 15.9 | 37.4 | |
| GRPOBase Model=Qwen2.5-Math-7B2026.02 | — | 12.6 | 25.6 | |
| GRPOBase Model=Qwen3-8B2026.02 | — | 20.2 | 38.6 | |
| Qwen3-1.7BModel Scale=1.7B, Method=Base2026.05 | — | — | 30 | |
| Qwen3-1.7B + DASDModel Scale=1.7B, Method=DASD2026.05 | — | — | 63.3 | |
| Qwen3-1.7B + GRPOModel Scale=1.7B, Method=GRPO2026.05 | — | — | 56.7 | |
| Qwen3-1.7B + HEPOModel Scale=1.7B, Method=HEPO2026.05 | — | — | 60 | |
| Qwen3-1.7B + OPSDModel Scale=1.7B, Method=OPSD2026.05 | — | — | 23.3 | |
| Qwen3-1.7B + RLSDModel Scale=1.7B, Method=RLSD2026.05 | — | — | 46.7 | |
| Qwen3-1.7B + SDPOModel Scale=1.7B, Method=SDPO2026.05 | — | — | 26.7 | |
| Qwen3-1.7B + SRPOModel Scale=1.7B, Method=SRPO2026.05 | — | — | 43.3 | |
| Qwen3-4BModel Scale=4B, Method=Base2026.05 | — | — | 43.3 | |
| Qwen3-4B + DASDModel Scale=4B, Method=DASD2026.05 | — | — | 83.3 | |
| Qwen3-4B + GRPOModel Scale=4B, Method=GRPO2026.05 | — | — | 86.7 | |
| Qwen3-4B + HEPOModel Scale=4B, Method=HEPO2026.05 | — | — | 76.7 | |
| Qwen3-4B + OPSDModel Scale=4B, Method=OPSD2026.05 | — | — | 40 | |
| Qwen3-4B + RLSDModel Scale=4B, Method=RLSD2026.05 | — | — | 76.7 | |
| Qwen3-4B + SDPOModel Scale=4B, Method=SDPO2026.05 | — | — | 36.7 | |
| Qwen3-4B + SRPOModel Scale=4B, Method=SRPO2026.05 | — | — | 60 | |
| Qwen3-8BModel Scale=8B, Method=Base2026.05 | — | — | 40 | |
| Qwen3-8B + DASDModel Scale=8B, Method=DASD2026.05 | — | — | 83.3 | |
| Qwen3-8B + GRPOModel Scale=8B, Method=GRPO2026.05 | — | — | 80 | |
| Qwen3-8B + HEPOModel Scale=8B, Method=HEPO2026.05 | — | — | 76.7 | |
| Qwen3-8B + OPSDModel Scale=8B, Method=OPSD2026.05 | — | — | 40 | |
| Qwen3-8B + RLSDModel Scale=8B, Method=RLSD2026.05 | — | — | 76.7 | |
| Qwen3-8B + SDPOModel Scale=8B, Method=SDPO2026.05 | — | — | 43.3 | |
| Qwen3-8B + SRPOModel Scale=8B, Method=SRPO2026.05 | — | — | 60 | |
| RPCBase Model=Qwen2.5-Math-7B2026.02 | — | 12.2 | 28.2 | |
| RPCBase Model=Qwen3-8B2026.02 | — | 20.1 | 37 | |
| URSBase Model=Qwen2.5-Math-7B2026.02 | — | 11.6 | 22.9 | |
| URSBase Model=Qwen3-8B2026.02 | — | 20.7 | 38.4 |