Safety Evaluation on HarmBench out-of-domain
0.62ASR (OM)No Attack
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| No AttackModel=Qwen2.5-7B-Instruct2026.06 | 0.62 | 8.44 | |
| BaseModel=Qwen2.5-7B-Instruct2026.06 | 0.62 | 8.44 | |
| Trajectory AlignmentModel=Llama-3.1-8B-Instruct2026.06 | 1.56 | 0 | |
| Circuit BreakersModel=Mistral-7B-Instruct2026.06 | 1.56 | 0.94 | |
| Trajectory AlignmentModel=Qwen2.5-7B-Instruct, Iteration=22026.06 | 2.19 | 1.87 | |
| Trajectory AlignmentModel=Mistral-7B-Instruct2026.06 | 2.5 | 4.06 | |
| Trajectory AlignmentModel=Mistral-7B-Instruct, Iteration=22026.06 | 2.5 | 0.94 | |
| Trajectory AlignmentModel=Llama-3.1-8B-Instruct, Iteration=12026.06 | 2.81 | 3.44 | |
| No AttackModel=Llama-3.1-8B-Instruct2026.06 | 3.12 | 4.19 | |
| BaseModel=Llama-3.1-8B-Instruct2026.06 | 3.12 | 4.19 | |
| SafeProbingModel=Qwen2.5-7B-Instruct2026.06 | 3.44 | 18.12 | |
| Trajectory AlignmentModel=Llama-3.1-8B-Instruct, Iteration=22026.06 | 3.44 | 3.75 | |
| Trajectory AlignmentModel=Qwen2.5-7B-Instruct, Iteration=12026.06 | 5.94 | 5.63 | |
| Trajectory AlignmentModel=Mistral-7B-Instruct, Iteration=12026.06 | 7.5 | 6.25 | |
| LATModel=Llama-3.1-8B-Instruct2026.06 | 10.31 | 3.12 | |
| SafeProbingModel=Mistral-7B-Instruct2026.06 | 12.19 | 20.26 | |
| Trajectory AlignmentModel=Qwen2.5-7B-Instruct2026.06 | 15.94 | 5 | |
| No AttackModel=Mistral-7B-Instruct2026.06 | 25 | 25.31 | |
| BaseModel=Mistral-7B-Instruct2026.06 | 25 | 25.31 | |
| Circuit BreakersModel=Llama-3.1-8B-Instruct2026.06 | 29.69 | 26.56 | |
| SafeProbingModel=Llama-3.1-8B-Instruct2026.06 | 37.5 | 53.44 | |
| Egida-DPOModel=Llama-3.1-8B-Instruct2026.06 | 40.31 | 44.37 | |
| Egida-DPOModel=Qwen2.5-7B-Instruct2026.06 | 46.88 | 53.44 | |
| Injection (no defense)Model=Mistral-7B-Instruct2026.06 | 47.5 | 43.44 | |
| Injection (no defense)Model=Qwen2.5-7B-Instruct2026.06 | 52.5 | 45 | |
| Injection (no defense)Model=Llama-3.1-8B-Instruct2026.06 | 61.25 | 48.13 |