Preference Alignment on PrefEval Explicit Preference (test)
77.7LLM-Evaluated ScoreDVTS w/ REAR
Evaluation Results
| Method | Links | |
|---|---|---|
| DVTS w/ REARMethod=DVTS, Reward Model=REAR, Base Model=Qwen2.5-7B-Instruct2026.06 | 77.7 | |
| BoN w/ REARMethod=BoN, Reward Model=REAR, Base Model=Qwen2.5-7B-Instruct, N=162026.06 | 74.1 | |
| BoN w/ GenRMMethod=BoN, Reward Model=GenRM, Base Model=Qwen2.5-7B-Instruct, N=162026.06 | 69 | |
| AmuletMethod=Amulet, Base Model=Qwen2.5-7B-Instruct2026.06 | 68.5 | |
| GreedyMethod=Greedy decoding, Base Model=Qwen2.5-7B-Instruct2026.06 | 67 | |
| LAMethod=LA, Base Model=Qwen2.5-7B-Instruct2026.06 | 64.2 |