LLM Alignment Evaluation on Qwen2.5-14B-Instruct High-Variance (Top 20%)
5.67Average Reward (μ)Base (Best-of-K)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Base (Best-of-K)Base Model=Qwen2.5-14B-Instruct, Strategy=Best-of-K2026.03 | 5.67 | — | -4.74 | 4.85 | |
| DARC-ϵBase Model=Qwen2.5-14B-Instruct2026.03 | 5.49 | — | -1.85 | 5.19 | |
| DARCBase Model=Qwen2.5-14B-Instruct2026.03 | 5.44 | — | -3.14 | 5.1 | |
| CVaR (Best-of-K)Base Model=Qwen2.5-14B-Instruct, Strategy=Best-of-K2026.03 | 5.41 | — | -4.59 | 4.99 | |
| DARC-τBase Model=Qwen2.5-14B-Instruct2026.03 | 5.41 | — | -2.91 | 5.12 | |
| 2nd-Moment (LCB)Base Model=Qwen2.5-14B-Instruct, Strategy=LCB2026.03 | 5.39 | — | -3.67 | 5.04 |