Mathematical Reasoning on GSM8K (test) (pass@1, pass@8)
75.52Pass@1Qwen3-1.7B-Base + DPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-1.7B-Base + DPOModel / Training=+ DPO, Scoring=Qwen3-30B2026.05 | 75.52 | 94.92 | |
| Qwen3-1.7B-Base + DPOModel / Training=+ DPO, Scoring=Inv. Entropy2026.05 | 71.17 | 94.01 | |
| Qwen3-1.7B-Base + σ-RRHFModel / Training=+ σ-RRHF, Scoring=Inv. Entropy, lambda=0.12026.05 | 70.5 | 93.8 | |
| Qwen3-1.7B-Base + σ-RRHFModel / Training=+ σ-RRHF, Scoring=Qwen3-30B, lambda=0.12026.05 | 70.29 | 93.48 | |
| Qwen3-1.7B-Base + DPOModel / Training=+ DPO, Scoring=Self-judge2026.05 | 69.34 | 93.71 | |
| Qwen3-1.7B-Base + SFTModel / Training=+ SFT2026.05 | 68.1 | 93.7 | |
| Qwen3-1.7B-BaseModel / Training=Qwen3-1.7B-Base2026.05 | 67.83 | 93.25 | |
| Qwen3-1.7B-Base + DPOModel / Training=+ DPO, Scoring=Random2026.05 | 67.74 | 93.56 | |
| Qwen3-1.7B-Base + σ-RRHFModel / Training=+ σ-RRHF, Scoring=Random, lambda=0.12026.05 | 64.57 | 94.09 | |
| Qwen3-1.7B-Base + σ-RRHFModel / Training=+ σ-RRHF, Scoring=Self-judge, lambda=0.12026.05 | 55.43 | 92.57 |