Role-playing Preference Alignment on Ping-Pong Bench (test)
3.07ScoreBoN w/ REAR
Evaluation Results
| Method | Links | |
|---|---|---|
| BoN w/ REARMethod=BoN, Reward Model=REAR, Base Model=Qwen2.5-7B-Instruct, N=162026.06 | 3.07 | |
| DVTS w/ REARMethod=DVTS, Reward Model=REAR, Base Model=Qwen2.5-7B-Instruct2026.06 | 3.03 | |
| BoN w/ GenRMMethod=BoN, Reward Model=GenRM, Base Model=Qwen2.5-7B-Instruct, N=162026.06 | 3.01 | |
| LAMethod=LA, Base Model=Qwen2.5-7B-Instruct2026.06 | 3.01 | |
| GreedyMethod=Greedy decoding, Base Model=Qwen2.5-7B-Instruct2026.06 | 2.97 | |
| AmuletMethod=Amulet, Base Model=Qwen2.5-7B-Instruct2026.06 | 2.87 |