Jailbreak Attack on Jailbreak Attack Suite
42GCG ASRNo Defense
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| No DefenseModel=Qwen2.5-7B-Instruct2026.01 | 42 | 28 | 70 | 43.2 | 84.8 | — | — | — | — | |
| Self-ReminderModel=Qwen2.5-7B-Instruct2026.01 | 34 | 15 | 59 | 42.4 | 25.8 | — | — | — | — | |
| PPLModel=Qwen2.5-7B-Instruct2026.01 | 33 | 23 | 70 | 38.6 | 82.8 | — | — | — | — | |
| RTBackbone=Llama2-7B-chat, Relative Compute=1x2026.03 | 21 | — | 12.2 | — | — | 2.2 | 10.6 | 11.1 | 10.2 | |
| Llama3-8B-instructBackbone=Llama3-8B-instruct, Relative Compute=0x2026.03 | 19.7 | — | 8.9 | — | — | 8.6 | 48.8 | 15.1 | 16.5 | |
| Llama2-7B-chatBackbone=Llama2-7B-chat, Relative Compute=0x2026.03 | 16.8 | — | 17.7 | — | — | 0 | 27.7 | 8.2 | 20.8 | |
| ICDModel=Qwen2.5-7B-Instruct2026.01 | 12 | 21 | 64 | 40.6 | 50.3 | — | — | — | — | |
| PPLModel=Llama-3-8B-Instruct2026.01 | 8 | 9 | 8 | 14.4 | 21.9 | — | — | — | — | |
| No DefenseModel=Llama-3-8B-Instruct2026.01 | 7 | 17 | 9 | 17.8 | 23.2 | — | — | — | — | |
| RTBackbone=Llama3-8B-instruct, Relative Compute=1x2026.03 | 3.9 | — | 14.3 | — | — | 0 | 13.5 | 1 | 3.3 | |
| Self-ExaminationModel=Llama-3-8B-Instruct2026.01 | 3 | 1 | 3 | 16 | 15.9 | — | — | — | — | |
| Self-ReminderModel=Llama-3-8B-Instruct2026.01 | 3 | 2 | 2 | 8 | 0 | — | — | — | — | |
| RT-EATBackbone=Llama2-7B-chat, Relative Compute=9x2026.03 | 1.9 | — | 3 | — | — | 0.2 | 4.3 | 0.7 | 0 | |
| ICDModel=Llama-3-8B-Instruct2026.01 | 1 | 14 | 1 | 6 | 1.3 | — | — | — | — | |
| SafeDecodingModel=Qwen2.5-7B-Instruct2026.01 | 1 | 2 | 14 | 13.8 | 1.3 | — | — | — | — | |
| RT-EAT-LATBackbone=Llama3-8B-instruct, Relative Compute=9x2026.03 | 0.9 | — | 3.3 | — | — | 0 | 6.8 | 0 | 0 | |
| R2D2Backbone=Llama2-7B-chat, Relative Compute=6558x2026.03 | 0.7 | — | 6.5 | — | — | 0 | 7.3 | 0 | 2.6 | |
| RT-EAT-LATBackbone=Llama2-7B-chat, Relative Compute=9x2026.03 | 0.7 | — | 2.5 | — | — | 0 | 2.9 | 0.6 | 0 | |
| SafeDecodingModel=Llama-3-8B-Instruct2026.01 | 0 | 0 | 0 | 0.2 | 0.7 | — | — | — | — | |
| SafeThinkerModel=Llama-3-8B-Instruct2026.01 | 0 | 0 | 0 | 0.2 | 0 | — | — | — | — | |
| Self-ExaminationModel=Qwen2.5-7B-Instruct2026.01 | 0 | 3 | 34 | 30.4 | 30.5 | — | — | — | — | |
| SafeThinkerModel=Qwen2.5-7B-Instruct2026.01 | 0 | 0 | 0 | 9.4 | 0.7 | — | — | — | — |