Preference Alignment on PrefEval Implicit Preference (test)
19.1ScoreDVTS w/ REAR
Evaluation Results
| Method | Links | |
|---|---|---|
| DVTS w/ REARMethod=DVTS, Reward Model=REAR, Base Model=Qwen2.5-7B-Instruct2026.06 | 19.1 | |
| BoN w/ REARMethod=BoN, Reward Model=REAR, Base Model=Qwen2.5-7B-Instruct, N=162026.06 | 16.2 | |
| AmuletMethod=Amulet, Base Model=Qwen2.5-7B-Instruct2026.06 | 13.1 | |
| BoN w/ GenRMMethod=BoN, Reward Model=GenRM, Base Model=Qwen2.5-7B-Instruct, N=162026.06 | 12.9 | |
| LAMethod=LA, Base Model=Qwen2.5-7B-Instruct2026.06 | 12.8 | |
| GreedyMethod=Greedy decoding, Base Model=Qwen2.5-7B-Instruct2026.06 | 12 |