Helpfulness-Safety Alignment on PKU-Alignment
5.497Useful ScoreGD2PO-Hard
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GD2PO-HardModel=Qwen2.5-7B-Instruct2026.06 | 5.497 | 6.891 | 6.194 | |
| GD2PO-HardModel=Llama3.2-3B-Instruct2026.06 | 5.49 | 6.893 | 6.192 | |
| GDPOModel=Llama3.2-3B-Instruct2026.06 | 5.483 | 6.874 | 6.179 | |
| GD2PO-SNRModel=Llama3.2-3B-Instruct, SNR threshold τ=0.82026.06 | 5.463 | 6.879 | 6.171 | |
| GRPOModel=Llama3.2-3B-Instruct2026.06 | 5.444 | 6.897 | 6.171 | |
| GD2PO-SNRModel=Qwen2.5-7B-Instruct, SNR threshold τ=0.82026.06 | 5.43 | 6.84 | 6.135 | |
| GDPOModel=Qwen2.5-7B-Instruct2026.06 | 5.426 | 6.845 | 6.136 | |
| GRPOModel=Qwen2.5-7B-Instruct2026.06 | 5.386 | 6.834 | 6.11 |