Preference Alignment on HH-RLHF (test)
87.4Win RateCW-IPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CW-IPOAlignment Objective=IPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 87.4 | — | — | |
| CW-IPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=IPO, Data Configuration=Confidence-Weighted2026.03 | 86.8 | — | — | |
| CW-rDPOAlignment Objective=rDPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 86.64 | — | — | |
| CW-rDPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=rDPO, Data Configuration=Confidence-Weighted2026.03 | 86.2 | — | — | |
| HumanAlignment Objective=IPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 84.16 | — | — | |
| Human BaselineModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=IPO, Data Configuration=Human2026.03 | 83.4 | — | — | |
| WS-DPOAlignment Objective=rDPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 83.01 | — | — | |
| WS-DPOAlignment Objective=DPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 82.26 | — | — | |
| WS-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=rDPO, Data Configuration=Weak model 30% annotations2026.03 | 82.2 | — | — | |
| HumanAlignment Objective=rDPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 82.07 | — | — | |
| WS-DPOAlignment Objective=IPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 81.88 | — | — | |
| CW-DPOAlignment Objective=DPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 81.5 | — | — | |
| WS-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=DPO, Data Configuration=Weak model 30% annotations2026.03 | 81.4 | — | — | |
| Human BaselineModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=rDPO, Data Configuration=Human2026.03 | 81.2 | — | — | |
| WS-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=IPO, Data Configuration=Weak model 30% annotations2026.03 | 81 | — | — | |
| CW-DPOModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=DPO, Data Configuration=Confidence-Weighted2026.03 | 80.6 | — | — | |
| HumanAlignment Objective=DPO, Model Setup=Qwen2.5-0.5B → Qwen2.5-14B2026.03 | 79.77 | — | — | |
| Human BaselineModel Pair=Qwen2.5-0.5B -> Qwen2.5-14B, Alignment Method=DPO, Data Configuration=Human2026.03 | 78.8 | — | — | |
| CW-IPOAlignment Objective=IPO, Model Setup=OPT-125M → OPT-13B2026.03 | 64.3 | — | — | |
| CW-IPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=IPO, Data Configuration=Confidence-Weighted2026.03 | 63.5 | — | — | |
| CW-rDPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=rDPO, Data Configuration=Confidence-Weighted2026.03 | 63 | — | — | |
| CW-rDPOAlignment Objective=rDPO, Model Setup=OPT-125M → OPT-13B2026.03 | 62.95 | — | — | |
| WS-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=IPO, Data Configuration=Weak model 30% annotations2026.03 | 62.8 | — | — | |
| WS-DPOAlignment Objective=IPO, Model Setup=OPT-125M → OPT-13B2026.03 | 61.79 | — | — | |
| CW-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=DPO, Data Configuration=Confidence-Weighted2026.03 | 61.3 | — | — | |
| CW-DPOAlignment Objective=DPO, Model Setup=OPT-125M → OPT-13B2026.03 | 60.79 | — | — | |
| HumanAlignment Objective=IPO, Model Setup=OPT-125M → OPT-13B2026.03 | 58.87 | — | — | |
| Human BaselineModel Pair=OPT-125M -> OPT-13B, Alignment Method=IPO, Data Configuration=Human2026.03 | 58.2 | — | — | |
| WS-DPOAlignment Objective=rDPO, Model Setup=OPT-125M → OPT-13B2026.03 | 57.63 | — | — | |
| WS-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=rDPO, Data Configuration=Weak model 30% annotations2026.03 | 57.6 | — | — | |
| Human BaselineModel Pair=OPT-125M -> OPT-13B, Alignment Method=DPO, Data Configuration=Human2026.03 | 56.9 | — | — | |
| HumanAlignment Objective=DPO, Model Setup=OPT-125M → OPT-13B2026.03 | 56.88 | — | — | |
| WS-DPOModel Pair=OPT-125M -> OPT-13B, Alignment Method=DPO, Data Configuration=Weak model 30% annotations2026.03 | 56.7 | — | — | |
| HumanAlignment Objective=rDPO, Model Setup=OPT-125M → OPT-13B2026.03 | 55.92 | — | — | |
| Human BaselineModel Pair=OPT-125M -> OPT-13B, Alignment Method=rDPO, Data Configuration=Human2026.03 | 55.9 | — | — | |
| WS-DPOAlignment Objective=DPO, Model Setup=OPT-125M → OPT-13B2026.03 | 55.01 | — | — | |
| FedBiscuitModel=Qwen 2, Number of Clients=102026.05 | — | 48.85 | 75.12 | |
| FedBiscuitModel=Qwen 2, Number of Clients=502026.05 | — | 44.21 | 71.45 | |
| FedBiscuitModel=Qwen 2, Number of Clients=1002026.05 | — | 42.33 | 69.42 | |
| FedBiscuitModel=Gemma-2B, Number of Clients=102026.05 | — | 51.65 | 82.45 | |
| FedBiscuitModel=Gemma-2B, Number of Clients=502026.05 | — | 46.21 | 78.12 | |
| FedBiscuitModel=Gemma-2B, Number of Clients=1002026.05 | — | 43.44 | 76.05 | |
| FedDPOModel=Qwen 2, Number of Clients=102026.05 | — | 48.12 | 77.34 | |
| FedDPOModel=Qwen 2, Number of Clients=502026.05 | — | 43.05 | 69.22 | |
| FedDPOModel=Qwen 2, Number of Clients=1002026.05 | — | 41.48 | 67.15 | |
| FedDPOModel=Gemma-2B, Number of Clients=102026.05 | — | 52.34 | 83.12 | |
| FedDPOModel=Gemma-2B, Number of Clients=502026.05 | — | 44.15 | 78.45 | |
| FedDPOModel=Gemma-2B, Number of Clients=1002026.05 | — | 41.22 | 75.33 | |
| FedVPA-GPModel=Qwen 2, Number of Clients=102026.05 | — | 66.45 | 89.21 | |
| FedVPA-GPModel=Qwen 2, Number of Clients=502026.05 | — | 58.32 | 84.05 | |
| FedVPA-GPModel=Qwen 2, Number of Clients=1002026.05 | — | 55.18 | 82.31 | |
| FedVPA-GPModel=Gemma-2B, Number of Clients=102026.05 | — | 73.21 | 96.34 | |
| FedVPA-GPModel=Gemma-2B, Number of Clients=502026.05 | — | 64.48 | 95.12 | |
| FedVPA-GPModel=Gemma-2B, Number of Clients=1002026.05 | — | 60.15 | 92.45 | |
| FedVPLModel=Qwen 2, Number of Clients=102026.05 | — | 62.24 | 84.56 | |
| FedVPLModel=Qwen 2, Number of Clients=502026.05 | — | 54.18 | 78.12 | |
| FedVPLModel=Qwen 2, Number of Clients=1002026.05 | — | 53.05 | 77.34 | |
| FedVPLModel=Gemma-2B, Number of Clients=102026.05 | — | 66.82 | 89.15 | |
| FedVPLModel=Gemma-2B, Number of Clients=502026.05 | — | 56.41 | 84.34 | |
| FedVPLModel=Gemma-2B, Number of Clients=1002026.05 | — | 53.25 | 80.42 |