Chain-sum arithmetic reasoning on Chain-sum (Dsaturated)
29.25Pass@1Qwen3-1.7B-Base + σ-RRHF
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-1.7B-Base + σ-RRHFModel / Training=+ σ-RRHF, Scoring=Inv. Entropy2026.05 | 29.25 | 61.75 | |
| Qwen3-1.7B-Base + σ-RRHFModel / Training=+ σ-RRHF, Scoring=Self-judge2026.05 | 21.16 | 54.75 | |
| Qwen3-1.7B-Base + SFTModel / Training=+ SFT2026.05 | 19.38 | 55.5 | |
| Qwen3-1.7B-Base + DPOModel / Training=+ DPO, Scoring=Inv. Entropy2026.05 | 16.28 | 58.75 | |
| Qwen3-1.7B-Base + DPOModel / Training=+ DPO, Scoring=Self-judge2026.05 | 13.97 | 52.75 | |
| Qwen3-1.7B-BaseModel / Training=Qwen3-1.7B-Base2026.05 | 10.63 | 51.5 |