Chain-sum arithmetic reasoning on Chain-sum D'strict
22.35Pass@1Qwen3-1.7B-Base + σ-RRHF
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-1.7B-Base + σ-RRHFModel / Training=+ σ-RRHF, Scoring=Inv. Entropy2026.05 | 22.35 | 59.25 | |
| Qwen3-1.7B-Base + σ-RRHFModel / Training=+ σ-RRHF, Scoring=Qwen3-30B2026.05 | 19.69 | 58 | |
| Qwen3-1.7B-Base + σ-RRHFModel / Training=+ σ-RRHF, Scoring=Self-judge2026.05 | 18.75 | 56 | |
| Qwen3-1.7B-Base + SFTModel / Training=+ SFT2026.05 | 14.62 | 51 | |
| Qwen3-1.7B-Base + DPOModel / Training=+ DPO, Scoring=Inv. Entropy2026.05 | 13.91 | 53.75 | |
| Qwen3-1.7B-Base + DPOModel / Training=+ DPO, Scoring=Self-judge2026.05 | 13.56 | 53 | |
| Qwen3-1.7B-BaseModel / Training=Qwen3-1.7B-Base2026.05 | 10.63 | 51.5 |