Safety Evaluation on StrongReject (test)
100StrongReject ScoreSFT-DPO + LoRA
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SFT-DPO + LoRABase Model=Qwen2.5-7B-Instruct, Alignment=SFT-DPO, Strategy=LoRA2026.02 | 100 | — | — | — | |
| SFT-DPO + LoRABase Model=Llama3.1-8B-Instruct, Alignment=SFT-DPO, Strategy=LoRA2026.02 | 99.93 | — | — | — | |
| SFT-DPOBase Model=Llama3.1-8B-Instruct, Alignment=SFT-DPO2026.02 | 99.87 | — | — | — | |
| SFTBase Model=Llama3.1-8B-Instruct, Alignment=SFT2026.02 | 97.27 | — | — | — | |
| TRIDENT-EDGEModel Backbone=GEMMA-7B, Alignment Dataset=TRIDENT-EDGE2025.05 | — | 32 | 2.16 | 18 | |
| UnalignedModel Backbone=GEMMA-7B, Alignment Status=Unaligned2025.05 | — | 77 | 3.82 | 66 | |
| WILDBREAKModel Backbone=GEMMA-7B, Alignment Dataset=WILDBREAK2025.05 | — | 60 | 3.12 | 49 |