Jailbreak Attack on HarmBench example-based Llama3 8B
5Attack Success RateCoA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CoAEvaluation Judge=GPT-Judge2025.08 | 5 | 1.5 | |
| AmpleGCGEvaluation Judge=GPT-Judge2025.08 | 6 | 1.39 | |
| I-GCGEvaluation Judge=GPT-Judge2025.08 | 7 | 1.5 | |
| PiFEvaluation Judge=GPT-Judge2025.08 | 8.5 | 2.37 | |
| PAIREvaluation Judge=GPT-Judge2025.08 | 15 | 2.93 | |
| ActorBreakerEvaluation Judge=GPT-Judge2025.08 | 23 | 3.94 | |
| CrescendoEvaluation Judge=GPT-Judge2025.08 | 25.5 | 3.3 | |
| GCGEvaluation Judge=GPT-Judge2025.08 | 35 | 1.44 | |
| ReNeLLMEvaluation Judge=GPT-Judge2025.08 | 43.5 | 4.12 | |
| GCG AttackHuman Readable=false, Zero-shot=true2025.12 | 44 | — | |
| AutoDAN TurboHuman Readable=true, Zero-shot=true2025.12 | 62 | — | |
| PAIR AttackHuman Readable=true, Zero-shot=true2025.12 | 66 | — | |
| AutoDANEvaluation Judge=GPT-Judge2025.08 | 68.5 | 4.47 | |
| AGILEGenerator LLM=Llama-3-8B-Instruct, p=5, NCand=25, Evaluation Judge=GPT-Judge2025.08 | 76 | 4.67 | |
| Adversarial ReasoningHuman Readable=true, Zero-shot=true2025.12 | 88 | — | |
| Adaptive AttackHuman Readable=false, Zero-shot=true2025.12 | 100 | — | |
| Jailbreak-ZeroHuman Readable=true, Zero-shot=true2025.12 | 100 | — |