Direct Preference Optimization on RLHFlow AlpacaEval 2.0
19.85LCWRDifficulty-Based Preference Data Selection
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Difficulty-Based Preference Data SelectionSelection Method=Ours2025.08 | 19.85 | 19.44 | |
| Full SetSelection Method=Full Set2025.08 | 18.74 | 17.93 | |
| RandomSelection Method=Random2025.08 | 18.57 | 18.13 | |
| ZIP†Selection Method=ZIP2025.08 | 18.34 | 18.06 | |
| SDPOSelection Method=SDPO2025.08 | 18.09 | 17.83 | |
| DiverseEvol†Selection Method=DiverseEvol2025.08 | 17.52 | 16.73 | |
| Tulu3-SFTSelection Method=SFT Baseline2025.08 | 2.57 | 2.16 |