LLM Alignment Evaluation on Qwen2.5-14B-Instruct Overall
6.31Reward (Avg μ)Base (Best-of-K)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Base (Best-of-K)Base Model=Qwen2.5-14B-Instruct, Strategy=Best-of-K2026.03 | 6.31 | 3.14 | 0.03 | 5.11 | |
| DARC-ϵBase Model=Qwen2.5-14B-Instruct2026.03 | 6.18 | 2.53 | 1.12 | 5.51 | |
| DARC-τBase Model=Qwen2.5-14B-Instruct2026.03 | 6.11 | 2.71 | 0.69 | 5.43 | |
| DARCBase Model=Qwen2.5-14B-Instruct2026.03 | 5.92 | 2.73 | 0.46 | 5.38 | |
| 2nd-Moment (LCB)Base Model=Qwen2.5-14B-Instruct, Strategy=LCB2026.03 | 5.81 | 2.83 | 0.15 | 5.23 | |
| CVaR (Best-of-K)Base Model=Qwen2.5-14B-Instruct, Strategy=Best-of-K2026.03 | 5.73 | 3.01 | -0.29 | 5.16 |