Mathematical Reasoning on OpenMathInstruct 2 (val)
65.19Pass@1GRPO Baseline
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GRPO BaselineBase model=Qwen2.5-Math-1.5B-Instruct, Training steps=2000, Hardware=8×H200 GPUs, Validation samples=20482026.03 | 65.19 | 77.49 | 82.28 | |
| HDPOTeacher type=frozen, λ=0.01, Base model=Qwen2.5-Math-1.5B-Instruct, Training steps=2000, Hardware=8×H200 GPUs, Divergence=JSD with global token-count normalization2026.03 | 65.19 | 78.12 | 82.18 | |
| HDPOTeacher type=drifting, λ=0.01, Base model=Qwen2.5-Math-1.5B-Instruct, Training steps=2000, Hardware=8×H200 GPUs, Divergence=JSD with global token-count normalization2026.03 | 65.14 | 78.61 | 82.71 | |
| HDPOTeacher type=frozen, λ=0.1, Base model=Qwen2.5-Math-1.5B-Instruct, Training steps=2000, Hardware=8×H200 GPUs, Divergence=JSD with global token-count normalization2026.03 | 63.04 | 78.12 | 83.98 | |
| HDPOTeacher type=drifting, λ=0.1, Base model=Qwen2.5-Math-1.5B-Instruct, Training steps=2000, Hardware=8×H200 GPUs, Divergence=JSD with global token-count normalization2026.03 | 62.94 | 78.56 | 83.64 |