Jailbreak Prompt Classification on AdvBench 520 prompts
93AccuracyKumar et al.
Evaluation Results
| Method | Links | |
|---|---|---|
| Kumar et al.Attack Vectors=Jailbreaking but not specific to phishing, Models=Llama 2, DistilBERT, Best Model=Llama 22025.07 | 93 |
| Method | Links | |
|---|---|---|
| Kumar et al.Attack Vectors=Jailbreaking but not specific to phishing, Models=Llama 2, DistilBERT, Best Model=Llama 22025.07 | 93 |