Operational decision-making on hand-crafted simulator
0.6Throughput ImprovementSFT + DPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SFT + DPOTraining iterations=2, Base Model=Qwen2.5-14B-Instruct2026.03 | 0.6 | 0.5 | |
| SFT + DPOTraining iterations=1, Base Model=Qwen2.5-14B-Instruct2026.03 | 0.3 | 0.1 | |
| DPO onlyTraining iterations=2, Base Model=Qwen2.5-14B-Instruct2026.03 | 0 | -0.2 | |
| Human DecisionsProtocol=replayed2026.03 | 0 | — | |
| DPO onlyTraining iterations=1, Base Model=Qwen2.5-14B-Instruct2026.03 | -0.4 | -0.7 | |
| SFT onlyBase Model=Qwen2.5-14B-Instruct2026.03 | -0.5 | -0.6 | |
| Self-RefineBase Model=Qwen 32B2026.03 | -2.1 | -2.4 |