Preference Alignment on PrefEval Implicit (test)
78.6Choice AccuracyDVTS w/ REAR
Evaluation Results
| Method | Links | |
|---|---|---|
| DVTS w/ REARMethod=DVTS, Reward Model=REAR, Base Model=Qwen2.5-7B-Instruct2026.06 | 78.6 | |
| BoN w/ REARMethod=BoN, Reward Model=REAR, Base Model=Qwen2.5-7B-Instruct, N=162026.06 | 78.2 | |
| LAMethod=LA, Base Model=Qwen2.5-7B-Instruct2026.06 | 78 | |
| BoN w/ GenRMMethod=BoN, Reward Model=GenRM, Base Model=Qwen2.5-7B-Instruct, N=162026.06 | 74.7 | |
| GreedyMethod=Greedy decoding, Base Model=Qwen2.5-7B-Instruct2026.06 | 71.5 | |
| AmuletMethod=Amulet, Base Model=Qwen2.5-7B-Instruct2026.06 | 70.4 |