Preference Alignment on HH and UF In-Domain (test)
68.4Win RateTPMM-DPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TPMM-DPOMerging Strategy=Learnable (Proposed), Iteration=32026.05 | 68.4 | 0.681 | |
| rDPO2026.05 | 65.9 | 0.664 | |
| TPMM-DPOMerging Strategy=Simple Average, Iteration=32026.05 | 63.7 | 0.629 | |
| TPMM-DPOMerging Strategy=Learnable (Proposed), Iteration=22026.05 | 62.9 | 0.671 | |
| TPMM-DPOMerging Strategy=Simple Average, Iteration=22026.05 | 62.4 | 0.643 | |
| sDPO2026.05 | 61.8 | 0.627 | |
| DPOIteration=22026.05 | 61.5 | 0.621 | |
| DPOIteration=32026.05 | 59.7 | 0.589 | |
| DPOIteration=12026.05 | 56.8 | 0.613 | |
| SFT2026.05 | 50 | 0.582 |