Preference Alignment on TL;DR (test)
68.8Win RateCW-rDPO
Evaluation Results
| Method | Links | |
|---|---|---|
| CW-rDPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=rDPO, Data Configuration=Confidence-Weighted2026.03 | 68.8 | |
| CW-rDPOAlignment Objective=rDPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 68.45 | |
| Human BaselineModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=rDPO, Data Configuration=Human2026.03 | 67 | |
| HumanAlignment Objective=rDPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 66.47 | |
| WS-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=rDPO, Data Configuration=Weak model 30% annotations2026.03 | 66.4 | |
| CW-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=DPO, Data Configuration=Confidence-Weighted2026.03 | 66 | |
| WS-DPOAlignment Objective=rDPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 65.87 | |
| CW-DPOAlignment Objective=DPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 65.67 | |
| WS-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=DPO, Data Configuration=Weak model 30% annotations2026.03 | 64.8 | |
| WS-DPOAlignment Objective=DPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 64.29 | |
| Human BaselineModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=DPO, Data Configuration=Human2026.03 | 64.2 | |
| CW-IPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=IPO, Data Configuration=Confidence-Weighted2026.03 | 64.2 | |
| HumanAlignment Objective=DPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 63.69 | |
| CW-IPOAlignment Objective=IPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 63.69 | |
| WS-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=IPO, Data Configuration=Weak model 30% annotations2026.03 | 62.8 | |
| CW-rDPOAlignment Objective=rDPO, Model Setup=OPT-125M → OPT-13B2026.03 | 62.12 | |
| WS-DPOAlignment Objective=IPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 62.1 | |
| Human BaselineModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=IPO, Data Configuration=Human2026.03 | 61.8 | |
| CW-rDPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=rDPO, Data Configuration=Confidence-Weighted2026.03 | 61.4 | |
| HumanAlignment Objective=IPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 61.31 | |
| HumanAlignment Objective=DPO, Model Setup=OPT-125M → OPT-13B2026.03 | 57.79 | |
| Human BaselineModel Pair=OPT-125M -> OPT-13B, Alignment Method=DPO, Data Configuration=Human2026.03 | 57 | |
| CW-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=DPO, Data Configuration=Confidence-Weighted2026.03 | 56.6 | |
| HumanAlignment Objective=rDPO, Model Setup=OPT-125M → OPT-13B2026.03 | 55.04 | |
| CW-IPOAlignment Objective=IPO, Model Setup=OPT-125M → OPT-13B2026.03 | 54.8 | |
| CW-IPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=IPO, Data Configuration=Confidence-Weighted2026.03 | 54.6 | |
| Human BaselineModel Pair=OPT-125M -> OPT-13B, Alignment Method=rDPO, Data Configuration=Human2026.03 | 54.2 | |
| CW-DPOAlignment Objective=DPO, Model Setup=OPT-125M → OPT-13B2026.03 | 53.85 | |
| WS-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=DPO, Data Configuration=Weak model 30% annotations2026.03 | 53.5 | |
| WS-DPOAlignment Objective=DPO, Model Setup=OPT-125M → OPT-13B2026.03 | 53.45 | |
| Human BaselineModel Pair=OPT-125M -> OPT-13B, Alignment Method=IPO, Data Configuration=Human2026.03 | 53.3 | |
| HumanAlignment Objective=IPO, Model Setup=OPT-125M → OPT-13B2026.03 | 51.54 | |
| WS-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=IPO, Data Configuration=Weak model 30% annotations2026.03 | 49.7 | |
| WS-DPOAlignment Objective=IPO, Model Setup=OPT-125M → OPT-13B2026.03 | 49.3 | |
| WS-DPOAlignment Objective=rDPO, Model Setup=OPT-125M → OPT-13B2026.03 | 48.14 | |
| WS-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=rDPO, Data Configuration=Weak model 30% annotations2026.03 | 47.7 |