Harmful Prompts
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Harmful Prompts Curated April 13, 2023
0Bad Bot Rate
61
Harmful Prompts
8.3Harmful Score
40
Harmful Prompts
2ASR (Raw)
15
Harmful prompts (evaluated on 3 LLMs and 4 guard LLMs)
3.23Mean Perplexity
10
Harmful Prompts Text-only baseline
0Text ASR
10
100 Harmful Prompts
55ASR (K=2)
9
Harmful Prompts model-averaged
4.8Model Averaged ASR (0.8)
8
Harmful Prompts
70.3ASR
8
Harmful Prompts 8-model (test)
92.4ASR
5
Harmful Prompts Victim: GPT-4o
4Runtime (H:MM)
5
Harmful Prompts SDXL
88Attack Success Rate (ASR)
4