Mathematical Reasoning on GMQ
35.3AccuracySFT (1Mixed)
Evaluation Results
| Method | Links | |
|---|---|---|
| SFT (1Mixed)Base Model=Qwen3-1.7B-Base, Training Strategy=SFT (1Mixed)2026.05 | 35.3 | |
| DGPOBase Model=Qwen3-1.7B-Base, Training Strategy=DGPO (Ours)2026.05 | 35 | |
| SFT (LIMO)Base Model=Qwen3-1.7B-Base, Training Strategy=SFT (LIMO)2026.05 | 34.8 | |
| SFT (3Mixed)Base Model=Qwen3-1.7B-Base, Training Strategy=SFT (3Mixed)2026.05 | 33.9 | |
| Qwen3-1.7B-BaseBase Model=Qwen3-1.7B-Base, Training Strategy=None2026.05 | 33.6 | |
| SFT (Reverse)Base Model=Qwen3-1.7B-Base, Training Strategy=SFT (Reverse)2026.05 | 33.6 | |
| DGPOBase Model=SFT (3Mixed), Training Strategy=DGPO (Ours)2026.05 | 33.6 | |
| Vanilla DPOBase Model=SFT (3Mixed), Training Strategy=Vanilla DPO2026.05 | 30.3 |