Preference Alignment on UFB (test)
81.05Win RateCW-DPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CW-DPOAlignment Objective=DPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 81.05 | — | |
| HumanAlignment Objective=IPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 80.41 | — | |
| CW-IPOAlignment Objective=IPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 80.17 | — | |
| WS-DPOAlignment Objective=DPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 78.12 | — | |
| HumanAlignment Objective=DPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 77.1 | — | |
| WS-DPOAlignment Objective=IPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 76.15 | — | |
| CW-rDPOAlignment Objective=rDPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 75.96 | — | |
| WS-DPOAlignment Objective=rDPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 74.03 | — | |
| HumanAlignment Objective=rDPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 73.12 | — | |
| CW-IPOAlignment Objective=IPO, Model Setup=OPT-125M → OPT-13B2026.03 | 65.17 | — | |
| CW-rDPOAlignment Objective=rDPO, Model Setup=OPT-125M → OPT-13B2026.03 | 63.95 | — | |
| WS-DPOAlignment Objective=DPO, Model Setup=OPT-125M → OPT-13B2026.03 | 63.72 | — | |
| CW-DPOAlignment Objective=DPO, Model Setup=OPT-125M → OPT-13B2026.03 | 63.61 | — | |
| WS-DPOAlignment Objective=rDPO, Model Setup=OPT-125M → OPT-13B2026.03 | 63.41 | — | |
| HumanAlignment Objective=IPO, Model Setup=OPT-125M → OPT-13B2026.03 | 62.63 | — | |
| HumanAlignment Objective=DPO, Model Setup=OPT-125M → OPT-13B2026.03 | 62.12 | — | |
| WS-DPOAlignment Objective=IPO, Model Setup=OPT-125M → OPT-13B2026.03 | 60.96 | — | |
| HumanAlignment Objective=rDPO, Model Setup=OPT-125M → OPT-13B2026.03 | 57.11 | — | |
| CW-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=DPO, Data Configuration=Confidence-Weighted2026.03 | — | 63.1 | |
| CW-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=DPO, Data Configuration=Confidence-Weighted2026.03 | — | 80.1 | |
| CW-IPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=IPO, Data Configuration=Confidence-Weighted2026.03 | — | 66.4 | |
| CW-IPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=IPO, Data Configuration=Confidence-Weighted2026.03 | — | 80.7 | |
| CW-rDPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=rDPO, Data Configuration=Confidence-Weighted2026.03 | — | 63.7 | |
| CW-rDPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=rDPO, Data Configuration=Confidence-Weighted2026.03 | — | 76.8 | |
| Human BaselineModel Pair=OPT-125M -> OPT-13B, Alignment Method=DPO, Data Configuration=Human2026.03 | — | 61.3 | |
| Human BaselineModel Pair=OPT-125M -> OPT-13B, Alignment Method=IPO, Data Configuration=Human2026.03 | — | 63.4 | |
| Human BaselineModel Pair=OPT-125M -> OPT-13B, Alignment Method=rDPO, Data Configuration=Human2026.03 | — | 58.9 | |
| Human BaselineModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=DPO, Data Configuration=Human2026.03 | — | 78.1 | |
| Human BaselineModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=IPO, Data Configuration=Human2026.03 | — | 78.5 | |
| Human BaselineModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=rDPO, Data Configuration=Human2026.03 | — | 72.4 | |
| WS-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=DPO, Data Configuration=Weak model 30% annotations2026.03 | — | 63.4 | |
| WS-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=IPO, Data Configuration=Weak model 30% annotations2026.03 | — | 61.3 | |
| WS-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=rDPO, Data Configuration=Weak model 30% annotations2026.03 | — | 61.2 | |
| WS-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=DPO, Data Configuration=Weak model 30% annotations2026.03 | — | 78.3 | |
| WS-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=IPO, Data Configuration=Weak model 30% annotations2026.03 | — | 77.2 | |
| WS-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=rDPO, Data Configuration=Weak model 30% annotations2026.03 | — | 75.1 |