Jailbreaking on JailbreakBench
2Attack Success Rate (ASR)Based
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| BasedTarget Model=Llama3.1-8B-Instruct2026.03 | 2 | — | — | — | |
| BaselineTarget Model=Llama-3-8B2026.05 | 3 | — | — | — | |
| BaselineTarget Model=Gemini-2.0-Flash2026.05 | 4 | — | — | — | |
| BaselineTarget Model=GPT-52026.05 | 5 | — | — | — | |
| BaselineTarget Model=Claude-3.7-Sonnet2026.05 | 6 | — | — | — | |
| BaselineTarget Model=GPT-5-Mini2026.05 | 6 | — | — | — | |
| ConVATarget Model=Qwen2.5-7B-Instruct2026.03 | 10 | — | — | — | |
| BaselineTarget Model=GPT-4o-Mini2026.05 | 10 | — | — | — | |
| BaselineTarget Model=Gemini-3-Flash2026.05 | 11 | — | — | — | |
| GCGTarget Model=Qwen2.5-7B-Instruct2026.03 | 12 | — | — | — | |
| BaselineTarget Model=Vicuna-13B2026.05 | 12 | — | — | — | |
| GCGTarget Model=Llama3.1-8B-Instruct2026.03 | 14 | — | — | — | |
| Quantum MechanicsTarget Model=Vicuna-13B2026.05 | 15 | — | — | — | |
| ConVATarget Model=Llama3.1-8B-Instruct2026.03 | 16 | — | — | — | |
| BasedTarget Model=Qwen2.5-7B-Instruct2026.03 | 18 | — | — | — | |
| Quantum MechanicsTarget Model=GPT-5-Mini2026.05 | 25 | — | — | — | |
| Quantum MechanicsTarget Model=GPT-52026.05 | 25 | — | — | — | |
| ConVATarget Model=Mistral-7B-Instruct2026.03 | 40 | — | — | — | |
| BasedTarget Model=Mistral-7B-Instruct2026.03 | 44 | — | — | — | |
| Set TheoryTarget Model=GPT-5-Mini2026.05 | 44 | — | — | — | |
| Quantum MechanicsTarget Model=Llama-3-8B2026.05 | 46 | — | — | — | |
| Formal LogicTarget Model=Vicuna-13B2026.05 | 47 | — | — | — | |
| Formal LogicTarget Model=GPT-52026.05 | 48 | — | — | — | |
| PAIRTarget Model=Llama3.1-8B-Instruct2026.03 | 52 | — | — | — | |
| Set TheoryTarget Model=GPT-52026.05 | 52 | — | — | — | |
| Quantum MechanicsTarget Model=GPT-4o-Mini2026.05 | 53 | — | — | — | |
| CAATarget Model=Llama3.1-8B-Instruct2026.03 | 54 | — | — | — | |
| Formal LogicTarget Model=GPT-5-Mini2026.05 | 57 | — | — | — | |
| CAATarget Model=Mistral-7B-Instruct2026.03 | 58 | — | — | — | |
| Set TheoryTarget Model=Vicuna-13B2026.05 | 59 | — | — | — | |
| Quantum MechanicsTarget Model=Claude-3.7-Sonnet2026.05 | 60 | — | — | — | |
| Quantum MechanicsTarget Model=Gemini-3-Flash2026.05 | 61 | — | — | — | |
| PAIRTarget Model=Qwen2.5-7B-Instruct2026.03 | 62 | — | — | — | |
| Quantum MechanicsTarget Model=Gemini-2.0-Flash2026.05 | 63 | — | — | — | |
| Set TheoryTarget Model=Claude-3.7-Sonnet2026.05 | 64 | — | — | — | |
| Formal LogicTarget Model=Claude-3.7-Sonnet2026.05 | 64 | — | — | — | |
| Formal LogicTarget Model=GPT-4o-Mini2026.05 | 65 | — | — | — | |
| GCGTarget Model=Mistral-7B-Instruct2026.03 | 66 | — | — | — | |
| PAIRTarget Model=Mistral-7B-Instruct2026.03 | 68 | — | — | — | |
| Formal LogicTarget Model=Gemini-2.0-Flash2026.05 | 68 | — | — | — | |
| Formal LogicTarget Model=Llama-3-8B2026.05 | 68 | — | — | — | |
| Set TheoryTarget Model=Gemini-3-Flash2026.05 | 69 | — | — | — | |
| SCAVTarget Model=Qwen2.5-7B-Instruct2026.03 | 70 | — | — | — | |
| Set TheoryTarget Model=GPT-4o-Mini2026.05 | 70 | — | — | — | |
| Formal LogicTarget Model=Gemini-3-Flash2026.05 | 70 | — | — | — | |
| Set TheoryTarget Model=Llama-3-8B2026.05 | 71 | — | — | — | |
| REATarget Model=Qwen2.5-7B-Instruct2026.03 | 76 | — | — | — | |
| Set TheoryTarget Model=Gemini-2.0-Flash2026.05 | 78 | — | — | — | |
| REATarget Model=Llama3.1-8B-Instruct2026.03 | 80 | — | — | — | |
| REATarget Model=Mistral-7B-Instruct2026.03 | 82 | — | — | — | |
| CAATarget Model=Qwen2.5-7B-Instruct2026.03 | 86 | — | — | — | |
| SCAVTarget Model=Llama3.1-8B-Instruct2026.03 | 90 | — | — | — | |
| SCAVTarget Model=Mistral-7B-Instruct2026.03 | 92 | — | — | — | |
| PAIRClassifier=JailbreakBench Classifier [35], Best of N responses=42025.12 | — | — | 4 | — | |
| PAIR with RT mutator LLMClassifier=JailbreakBench Classifier [35], Best of N responses=42025.12 | — | 1 | 1 | — | |
| PAIR with RT mutator LLMClassifier=Llama Guard (JBB Behaviours), Best of N responses=42025.12 | — | 14 | 11 | — | |
| PAIR with RT mutator LLMDescription=Prompt diversity score2025.12 | — | — | — | 0.74 | |
| Rainbow TeamingClassifier=JailbreakBench Classifier [35], Best of N responses=42025.12 | — | 8 | 7 | — | |
| Rainbow TeamingClassifier=Llama Guard (JBB Behaviours), Best of N responses=42025.12 | — | 66 | 41 | — | |
| Rainbow TeamingDescription=Prompt diversity score2025.12 | — | — | — | 0.51 |