Reward Modeling Evaluation on UltraFeedback (test)
-3.12ScoreDPO+Filter
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DPO+FilterPη Type=12025.10 | -3.12 | 67 | |
| DPOPη Type=12025.10 | -3.59 | 63 | |
| DPO+FilterPη Type=22025.10 | -3.75 | 64 | |
| DPOPη Type=22025.10 | -4.42 | 58 | |
| BasePη Type=-2025.10 | -5.47 | 50 |