Adversarial Alignment Robustness on Handcrafted jailbreak prompts
99.3BARVicuna-7B-chat-HF
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Vicuna-7B-chat-HFAlignment Method=Original LLM2023.09 | 99.3 | 98.7 | — | |
| GPT-3.5-turbo-0613Alignment Method=Original LLM2023.09 | 99.3 | 82 | — | |
| GPT-3.5-turbo-0613Alignment Method=RA-LLM2023.09 | 99.3 | 8 | 74 | |
| Vicuna-7B-chat-HFAlignment Method=RA-LLM2023.09 | 98.7 | 12 | 86.7 | |
| Guanaco-7B-HFAlignment Method=Original LLM2023.09 | 95.3 | 94.7 | — | |
| Guanaco-7B-HFAlignment Method=RA-LLM2023.09 | 92 | 9.3 | 85.4 |