Helpfulness-Safety Alignment on HH-RLHF
4.674Useful ScoreGD2PO-Hard
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GD2PO-HardModel=Llama3.2-3B-Instruct2026.06 | 4.674 | 5.863 | 5.269 | |
| GD2PO-SNRModel=Llama3.2-3B-Instruct, SNR threshold τ=0.82026.06 | 4.62 | 5.843 | 5.232 | |
| GDPOModel=Llama3.2-3B-Instruct2026.06 | 4.594 | 5.8 | 5.197 | |
| GRPOModel=Llama3.2-3B-Instruct2026.06 | 4.559 | 5.849 | 5.204 | |
| GD2PO-HardModel=Qwen2.5-7B-Instruct2026.06 | 4.502 | 5.708 | 5.105 | |
| GD2PO-SNRModel=Qwen2.5-7B-Instruct, SNR threshold τ=0.82026.06 | 4.483 | 5.659 | 5.071 | |
| GDPOModel=Qwen2.5-7B-Instruct2026.06 | 4.408 | 5.619 | 5.014 | |
| GRPOModel=Qwen2.5-7B-Instruct2026.06 | 4.237 | 5.469 | 4.853 |