Jailbreak Defense Performance on Jailbreak Attack Dataset
96.2DSRVanilla
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VanillaModel=Vicuna-13B2023.11 | 96.2 | — | — | — | 95.7 | 95.8 | 66.9 | 50.5 | 37.5 | 5 | 0 | 64.5 | |
| VanillaModel=Vicuna-33B2023.11 | 96.2 | — | — | — | 100 | 96.7 | 70.6 | 51 | 52.5 | 15 | 5 | 68.2 | |
| VanillaModel=Vicuna-7B2023.11 | 94.4 | — | — | — | 87.1 | 75 | 55.6 | 44.5 | 37.5 | 7.5 | 1.7 | 57.8 | |
| VanillaModel=ChatGPT2023.11 | 93.8 | — | — | — | 87.1 | 75 | 56.9 | 79 | 41.2 | 21.2 | 5 | 66.4 | |
| CAVGANModel=Qwen2.5-7B2025.07 | 91.12 | 91.4 | 88.14 | 10.06 | — | — | — | — | — | — | — | — | |
| Self-ReminderModel=Vicuna-33B2023.11 | 80.6 | — | — | — | 100 | 92.5 | 43.1 | 49 | 7.5 | 3.8 | 11.7 | 56.3 | |
| RA-LLMModel=Qwen2.5-7B2025.07 | 78.6 | 85.8 | 71.42 | 12.45 | — | — | — | — | — | — | — | — | |
| CAVGANModel=Llama3.1-8B2025.07 | 77.22 | 93.6 | 74.31 | 6.02 | — | — | — | — | — | — | — | — | |
| CAVGANModel=Mistral-8B2025.07 | 76.37 | 91.06 | 70.97 | 8.57 | — | — | — | — | — | — | — | — | |
| Goal PrioritizationModel=Vicuna-7B2023.11 | 75.6 | — | — | — | 63.6 | 59.2 | 18.8 | 17.5 | 2.5 | 2.5 | 0 | 35 | |
| Self-ReminderModel=Vicuna-7B2023.11 | 73.8 | — | — | — | 87.9 | 70.8 | 24.4 | 34.5 | 2.5 | 1.2 | 0 | 43.7 | |
| RA-LLMModel=Llama3.1-8B2025.07 | 73.78 | 92.8 | 70.42 | 6.83 | — | — | — | — | — | — | — | — | |
| RA-LLMModel=Mistral-8B2025.07 | 71.18 | 89.2 | 64.59 | 10.44 | — | — | — | — | — | — | — | — | |
| VanillaModel=GPT-42023.11 | 70.6 | — | — | — | 20 | 75.8 | 36.9 | 62.5 | 42.5 | 21.2 | 26.7 | 48.3 | |
| Self-ReminderModel=Vicuna-13B2023.11 | 68.8 | — | — | — | 97.1 | 92.5 | 25.6 | 50.5 | 18.8 | 1.2 | 1.7 | 51.6 | |
| Smooth-llmModel=Qwen2.5-7B2025.07 | 54.22 | 75.77 | 38.86 | 22.68 | — | — | — | — | — | — | — | — | |
| Smooth-llmModel=Mistral-8B2025.07 | 52.07 | 81.85 | 41.12 | 17.82 | — | — | — | — | — | — | — | — | |
| Smooth-llmModel=Llama3.1-8B2025.07 | 48.97 | 81.03 | 42.44 | 18.64 | — | — | — | — | — | — | — | — | |
| Self-ReminderModel=ChatGPT2023.11 | 37.5 | — | — | — | 65.7 | 27.5 | 31.2 | 18 | 6.2 | 3.8 | 3.3 | 28.1 | |
| Goal PrioritizationModel=Vicuna-13B2023.11 | 36.9 | — | — | — | 47.9 | 34.2 | 8.8 | 10 | 5 | 2.5 | 1.7 | 20.8 | |
| Goal PrioritizationModel=Vicuna-33B2023.11 | 26.9 | — | — | — | 46.4 | 27.5 | 8.8 | 17 | 1.2 | 2.5 | 0 | 19.2 | |
| OriginalModel=Qwen2.5-7B2025.07 | 25.12 | 98 | — | — | — | — | — | — | — | — | — | — | |
| OriginalModel=Mistral-8B2025.07 | 18.59 | 99.6 | — | — | — | — | — | — | — | — | — | — | |
| OriginalModel=Llama3.1-8B2025.07 | 11.34 | 99.6 | — | — | — | — | — | — | — | — | — | — | |
| VanillaModel=Llama2-13B-Chat2023.11 | 11 | — | — | — | 15 | 16.7 | 5 | 65.5 | 5 | 1.2 | 0 | 21 | |
| VanillaModel=Llama2-7B-Chat2023.11 | 4 | — | — | — | 15.8 | 21.7 | 3.1 | 73.5 | 3.8 | 1.2 | 0 | 22.2 | |
| Goal PrioritizationModel=ChatGPT2023.11 | 2.5 | — | — | — | 5 | 1.7 | 3.8 | 5.5 | 5 | 2.5 | 0 | 3.6 | |
| Self-ReminderModel=GPT-42023.11 | 2.5 | — | — | — | 5 | 26.7 | 6.2 | 4.5 | 6.2 | 5 | 1.7 | 7.2 | |
| Goal PrioritizationModel=GPT-42023.11 | 1.9 | — | — | — | 0 | 1.7 | 10.6 | 3.5 | 1.2 | 1.2 | 0 | 3.1 | |
| Goal PrioritizationModel=Llama2-13B-Chat2023.11 | 1.9 | — | — | — | 2.9 | 0.8 | 0 | 8 | 1.2 | 0 | 0 | 2.5 | |
| Self-ReminderModel=Llama2-13B-Chat2023.11 | 1.2 | — | — | — | 5.7 | 0 | 1.9 | 17 | 1.2 | 0 | 0 | 4.8 | |
| Self-ReminderModel=Llama2-7B-Chat2023.11 | 0.6 | — | — | — | 4.3 | 5 | 1.2 | 44 | 0 | 0 | 0 | 10.3 | |
| Goal PrioritizationModel=Llama2-7B-Chat2023.11 | 0.6 | — | — | — | 5 | 3.3 | 1.9 | 9 | 0 | 0 | 5 | 3.6 |