Safety Evaluation on JailbreakBench (out-of-domain)
73ASR (OM)Injection (no defense)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Injection (no defense)Model=Llama-3.1-8B-Instruct2026.06 | 73 | 66 | |
| Injection (no defense)Model=Mistral-7B-Instruct2026.06 | 72 | 64 | |
| Injection (no defense)Model=Qwen2.5-7B-Instruct2026.06 | 72 | 67 | |
| Egida-DPOModel=Qwen2.5-7B-Instruct2026.06 | 66 | 74 | |
| Egida-DPOModel=Llama-3.1-8B-Instruct2026.06 | 49 | 53 | |
| SafeProbingModel=Llama-3.1-8B-Instruct2026.06 | 49 | 55 | |
| No AttackModel=Mistral-7B-Instruct2026.06 | 32 | 31 | |
| BaseModel=Mistral-7B-Instruct2026.06 | 32 | 31 | |
| Circuit BreakersModel=Llama-3.1-8B-Instruct2026.06 | 29 | 33 | |
| SafeProbingModel=Mistral-7B-Instruct2026.06 | 15 | 20 | |
| Trajectory AlignmentModel=Qwen2.5-7B-Instruct2026.06 | 8 | 5 | |
| No AttackModel=Qwen2.5-7B-Instruct2026.06 | 5 | 0 | |
| Trajectory AlignmentModel=Mistral-7B-Instruct, Iteration=12026.06 | 5 | 5 | |
| BaseModel=Qwen2.5-7B-Instruct2026.06 | 5 | 0 | |
| LATModel=Llama-3.1-8B-Instruct2026.06 | 4 | 0 | |
| Trajectory AlignmentModel=Qwen2.5-7B-Instruct, Iteration=12026.06 | 4 | 0 | |
| Trajectory AlignmentModel=Llama-3.1-8B-Instruct2026.06 | 3 | 0 | |
| SafeProbingModel=Qwen2.5-7B-Instruct2026.06 | 3 | 10 | |
| Trajectory AlignmentModel=Mistral-7B-Instruct, Iteration=22026.06 | 1 | 0 | |
| Trajectory AlignmentModel=Llama-3.1-8B-Instruct, Iteration=12026.06 | 1 | 0 | |
| Trajectory AlignmentModel=Qwen2.5-7B-Instruct, Iteration=22026.06 | 0.74 | 0 | |
| No AttackModel=Llama-3.1-8B-Instruct2026.06 | 0 | 0 | |
| Circuit BreakersModel=Mistral-7B-Instruct2026.06 | 0 | 2 | |
| Trajectory AlignmentModel=Mistral-7B-Instruct2026.06 | 0 | 2 | |
| BaseModel=Llama-3.1-8B-Instruct2026.06 | 0 | 0 | |
| Trajectory AlignmentModel=Llama-3.1-8B-Instruct, Iteration=22026.06 | 0 | 0 |