Reward Modeling on RewardBench RLHFlow source Chat Chat-Hard Safety Reasoning Total 1.0 (train)
80.62Chat ScoreDifficulty-Based Preference Data Selection
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Difficulty-Based Preference Data SelectionSelection Strategy=Ours2025.08 | 80.62 | 70.98 | 82.19 | 69.85 | 75.24 | |
| ZIPSelection Strategy=ZIP, Adapted from=IFT-oriented data selection2025.08 | 79.83 | 71.42 | 80.93 | 72.65 | 76.14 | |
| SDPOSelection Strategy=SDPO2025.08 | 79.61 | 70.9 | 79.42 | 69.57 | 75.15 | |
| DiverseEvolSelection Strategy=DiverseEvol, Adapted from=IFT-oriented data selection2025.08 | 78.55 | 70.24 | 79.56 | 70.38 | 74.93 | |
| Full SetSelection Strategy=Full Set2025.08 | 72.91 | 71.27 | 80.81 | 77.23 | 75.62 | |
| RandomSelection Strategy=Random2025.08 | 71.52 | 69.38 | 79.14 | 75.58 | 73.92 |