Safety Evaluation on AdvBench (in-domain)
0ASR (OM)Circuit Breakers
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Circuit BreakersModel=Mistral-7B-Instruct2026.06 | 0 | 0.19 | |
| Trajectory AlignmentModel=Mistral-7B-Instruct2026.06 | 0 | 0.96 | |
| Trajectory AlignmentModel=Mistral-7B-Instruct, Iteration=22026.06 | 0 | 0 | |
| No AttackModel=Llama-3.1-8B-Instruct2026.06 | 0.19 | 0 | |
| BaseModel=Llama-3.1-8B-Instruct2026.06 | 0.19 | 0 | |
| Trajectory AlignmentModel=Qwen2.5-7B-Instruct, Iteration=22026.06 | 0.96 | 0.19 | |
| Trajectory AlignmentModel=Llama-3.1-8B-Instruct, Iteration=12026.06 | 1.35 | 0 | |
| No AttackModel=Qwen2.5-7B-Instruct2026.06 | 1.54 | 0.96 | |
| BaseModel=Qwen2.5-7B-Instruct2026.06 | 1.54 | 0.96 | |
| Trajectory AlignmentModel=Llama-3.1-8B-Instruct, Iteration=22026.06 | 2.88 | 0 | |
| SafeProbingModel=Qwen2.5-7B-Instruct2026.06 | 3.27 | 4.81 | |
| Trajectory AlignmentModel=Qwen2.5-7B-Instruct, Iteration=12026.06 | 3.85 | 0.38 | |
| Trajectory AlignmentModel=Llama-3.1-8B-Instruct2026.06 | 4.42 | 0.19 | |
| Trajectory AlignmentModel=Mistral-7B-Instruct, Iteration=12026.06 | 4.42 | 2.12 | |
| LATModel=Llama-3.1-8B-Instruct2026.06 | 5.19 | 0 | |
| Trajectory AlignmentModel=Qwen2.5-7B-Instruct2026.06 | 11.73 | 5.19 | |
| SafeProbingModel=Mistral-7B-Instruct2026.06 | 14.23 | 14.04 | |
| No AttackModel=Mistral-7B-Instruct2026.06 | 38.27 | 31.54 | |
| BaseModel=Mistral-7B-Instruct2026.06 | 38.27 | 31.54 | |
| Circuit BreakersModel=Llama-3.1-8B-Instruct2026.06 | 39.81 | 38.27 | |
| SafeProbingModel=Llama-3.1-8B-Instruct2026.06 | 69.04 | 72.5 | |
| Egida-DPOModel=Llama-3.1-8B-Instruct2026.06 | 82.12 | 80.38 | |
| Injection (no defense)Model=Mistral-7B-Instruct2026.06 | 85.77 | 82.49 | |
| Injection (no defense)Model=Qwen2.5-7B-Instruct2026.06 | 88.08 | 87.73 | |
| Egida-DPOModel=Qwen2.5-7B-Instruct2026.06 | 88.65 | 85.77 | |
| Injection (no defense)Model=Llama-3.1-8B-Instruct2026.06 | 92.12 | 89.7 |