Helpfulness-Safety Alignment on Overall
5.808Average ScoreGD2PO-Hard
Evaluation Results
| Method | Links | |
|---|---|---|
| GD2PO-HardModel=Llama3.2-3B-Instruct2026.06 | 5.808 | |
| GD2PO-SNRModel=Llama3.2-3B-Instruct, SNR threshold τ=0.82026.06 | 5.779 | |
| GDPOModel=Llama3.2-3B-Instruct2026.06 | 5.76 | |
| GRPOModel=Llama3.2-3B-Instruct2026.06 | 5.756 | |
| GD2PO-HardModel=Qwen2.5-7B-Instruct2026.06 | 5.703 | |
| GD2PO-SNRModel=Qwen2.5-7B-Instruct, SNR threshold τ=0.82026.06 | 5.667 | |
| GDPOModel=Qwen2.5-7B-Instruct2026.06 | 5.6 | |
| GRPOModel=Qwen2.5-7B-Instruct2026.06 | 5.425 |