Jailbreak Defense on JailbreakBench (JBB)
36.11FNRTRLM-FO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TRLM-FOThresh=2, Training Protocol=Pretrained (PT)2024.12 | 36.11 | — | |
| TRLM-BaThresh=2, Training Protocol=Pretrained (PT)2024.12 | 52.78 | — | |
| TRLM-FOThresh=4, Training Protocol=Pretrained (PT)2024.12 | 55.56 | — | |
| TRLM-FOThresh=2, Training Protocol=Instruction-finetuned (IT)2024.12 | 55.56 | — | |
| TRLM-BaThresh=2, Training Protocol=Instruction-finetuned (IT)2024.12 | 59.72 | — | |
| TRLM-BaThresh=4, Training Protocol=Pretrained (PT)2024.12 | 65.28 | — | |
| TRLM-BaThresh=6, Training Protocol=Pretrained (PT)2024.12 | 69.44 | — | |
| TRLM-FOThresh=6, Training Protocol=Pretrained (PT)2024.12 | 70.83 | — | |
| TRLM-BaThresh=4, Training Protocol=Instruction-finetuned (IT)2024.12 | 70.83 | — | |
| TRLM-FOThresh=4, Training Protocol=Instruction-finetuned (IT)2024.12 | 72.22 | — | |
| TRLM-BaThresh=6, Training Protocol=Instruction-finetuned (IT)2024.12 | 79.17 | — | |
| TRLM-FOThresh=6, Training Protocol=Instruction-finetuned (IT)2024.12 | 81.94 | — | |
| DGR SFTAlignment Dataset=DirectRefusal, Backbone=s1.1-7B2026.02 | — | 99 | |
| DGR SFTAlignment Dataset=STAR-1, Backbone=s1.1-7B2026.02 | — | 99 | |
| DGR SFTAlignment Dataset=R1-ACT, Backbone=s1.1-7B2026.02 | — | 97 | |
| Qwen2.5-7B-Instruct2026.02 | — | 97 | |
| s1.1-7B2026.02 | — | 28 | |
| Vanilla SFTAlignment Dataset=DirectRefusal, Backbone=s1.1-7B2026.02 | — | 100 | |
| Vanilla SFTAlignment Dataset=STAR-1, Backbone=s1.1-7B2026.02 | — | 100 | |
| Vanilla SFTAlignment Dataset=R1-ACT, Backbone=s1.1-7B2026.02 | — | 97 |