Preference Alignment on HH and UF Out-of-Domain (test)
58.7OOD Win RateTPMM-DPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TPMM-DPOMerging Strategy=Learnable (Proposed), Iteration=32026.05 | 58.7 | 0.287 | |
| rDPO2026.05 | 57.9 | 0.271 | |
| TPMM-DPOMerging Strategy=Learnable (Proposed), Iteration=22026.05 | 57.6 | 0.242 | |
| DPOIteration=22026.05 | 57.4 | 0.238 | |
| TPMM-DPOMerging Strategy=Simple Average, Iteration=32026.05 | 57.2 | 0.266 | |
| TPMM-DPOMerging Strategy=Simple Average, Iteration=22026.05 | 57 | 0.234 | |
| sDPO2026.05 | 55.3 | 0.226 | |
| DPOIteration=12026.05 | 54.2 | 0.214 | |
| DPOIteration=32026.05 | 51.1 | 0.201 | |
| SFT2026.05 | 50 | 0.105 |