Mathematical Reasoning on OlympiadBench (Pass@1, Avg tokens)
57.1Pass@1 AccuracyNFPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| NFPOBase Model=Qwen3-8B-Base, Algorithm=NFPO2026.05 | 57.1 | — | |
| DPPOBase Model=Qwen3-8B-Base, Algorithm=DPPO2026.05 | 56.5 | — | |
| GRPO-LEADTraining Step=200, Pre-collapse peak=true2026.06 | 56.4 | 8,780 | |
| GRPO-acc*Training Step=1,000, Accuracy-only GRPO=true2026.06 | 56.2 | 7,982 | |
| BaseMethod variant=Baseline2026.06 | 55.3 | 9,579 | |
| ACOERTraining Step=1,2002026.06 | 55.3 | 5,177 | |
| ReCutTraining Step=400, Pre-collapse peak=true2026.06 | 55 | 5,888 | |
| GRPO+LPTraining Step=400, Pre-collapse peak=true, Length Penalty (LP)=true2026.06 | 54.2 | 7,150 | |
| GRPOBase Model=Qwen3-8B-Base, Algorithm=GRPO2026.05 | 52.7 | — | |
| NFPOBase Model=Qwen3-1.7B-Base, Algorithm=NFPO2026.05 | 36 | — | |
| PROGRS-4rollout_budget=42026.02 | 35.74 | 990.17 | |
| PROGRS-8rollout_budget=82026.02 | 35.06 | 958.38 | |
| DPPOBase Model=Qwen3-1.7B-Base, Algorithm=DPPO2026.05 | 34.1 | — | |
| PROGRS-8 (αcoh = 0)rollout_budget=8, coherence_penalty=disabled2026.02 | 33.52 | 934.8 | |
| β=0 (s42)Training Step=1,200, Beta coefficient (β)=02026.06 | 33.2 | 4,767 | |
| GRPO+LPTraining Step=1,200, Length Penalty (LP)=true2026.06 | 32.5 | 1,222 | |
| DAPO-16rollout_budget=162026.02 | 32.03 | 910.65 | |
| GRPOBase Model=Qwen3-1.7B-Base, Algorithm=GRPO2026.05 | 31.5 | — | |
| DAPO-8rollout_budget=82026.02 | 31.34 | 944.48 | |
| ReCutTraining Step=1,2002026.06 | 31.3 | 2,401 | |
| PROGRS-8 (No Centering)rollout_budget=8, centering=none2026.02 | 30.77 | 975.54 | |
| GRPO-LEADTraining Step=1,2002026.06 | 30.1 | 2,358 | |
| Qwen3-8B-BaseBase Model=Qwen3-8B-Base, Algorithm=Base2026.05 | 29.2 | — | |
| Qwen3-1.7B-BaseBase Model=Qwen3-1.7B-Base, Algorithm=Base2026.05 | 18.4 | — |