Helpfulness-Safety Alignment on Alpaca
5.633Useful ScoreGD2PO-Hard
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GD2PO-HardModel=Llama3.2-3B-Instruct2026.06 | 5.633 | 6.296 | 5.965 | |
| GDPOModel=Llama3.2-3B-Instruct2026.06 | 5.596 | 6.212 | 5.904 | |
| GD2PO-SNRModel=Llama3.2-3B-Instruct, SNR threshold τ=0.82026.06 | 5.577 | 6.29 | 5.934 | |
| GD2PO-SNRModel=Qwen2.5-7B-Instruct, SNR threshold τ=0.82026.06 | 5.527 | 6.061 | 5.794 | |
| GRPOModel=Llama3.2-3B-Instruct2026.06 | 5.514 | 6.272 | 5.893 | |
| GD2PO-HardModel=Qwen2.5-7B-Instruct2026.06 | 5.493 | 6.129 | 5.811 | |
| GDPOModel=Qwen2.5-7B-Instruct2026.06 | 5.36 | 5.944 | 5.652 | |
| GRPOModel=Qwen2.5-7B-Instruct2026.06 | 4.976 | 5.647 | 5.312 |