Reasoning on LMGH
12.9Accuracy (LMGH)SFT (Reverse)
Evaluation Results
| Method | Links | |
|---|---|---|
| SFT (Reverse)Base Model=Qwen3-1.7B-Base, Training Strategy=SFT (Reverse)2026.05 | 12.9 | |
| DGPOBase Model=SFT (3Mixed), Training Strategy=DGPO (Ours)2026.05 | 12.1 | |
| SFT (LIMO)Base Model=Qwen3-1.7B-Base, Training Strategy=SFT (LIMO)2026.05 | 11.4 | |
| SFT (3Mixed)Base Model=Qwen3-1.7B-Base, Training Strategy=SFT (3Mixed)2026.05 | 11.3 | |
| SFT (1Mixed)Base Model=Qwen3-1.7B-Base, Training Strategy=SFT (1Mixed)2026.05 | 9.9 | |
| Vanilla DPOBase Model=SFT (3Mixed), Training Strategy=Vanilla DPO2026.05 | 4.5 | |
| DGPOBase Model=Qwen3-1.7B-Base, Training Strategy=DGPO (Ours)2026.05 | 3 | |
| Qwen3-1.7B-BaseBase Model=Qwen3-1.7B-Base, Training Strategy=None2026.05 | 2.3 |