Jailbreak Success Rate (GPT-4o/5 Judges) and Query Cost on AdvBench
99.1ASR (GPT-4o)HMNS
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HMNSModel=Yi-1.5-72B-Chat2026.04 | 99.1 | 95.2 | 1.8 | |
| HMNSModel=Mistral-7B-Instruct-v0.22026.04 | 97.4 | 92.5 | 2 | |
| ArrAttackModel=Yi-1.5-72B-Chat2026.04 | 93.4 | 89.2 | 7.3 | |
| HMNSModel=Qwen2.5-14B-Instruct2026.04 | 92.9 | 86.8 | 2 | |
| ArrAttackModel=Mistral-7B-Instruct-v0.22026.04 | 91.1 | 86.2 | 7.6 | |
| ArrAttackModel=Qwen2.5-14B-Instruct2026.04 | 86.8 | 80.7 | 8.3 | |
| AutoDANModel=Yi-1.5-72B-Chat2026.04 | 73.9 | 67.6 | 12.3 | |
| AutoDANModel=Mistral-7B-Instruct-v0.22026.04 | 70.8 | 64.6 | 12.7 | |
| AutoDANModel=Qwen2.5-14B-Instruct2026.04 | 64.9 | 58.2 | 13.8 | |
| FITDModel=Yi-1.5-72B-Chat2026.04 | 45.9 | 40.1 | 15.6 | |
| FITDModel=Mistral-7B-Instruct-v0.22026.04 | 42 | 36.5 | 16.4 | |
| FITDModel=Qwen2.5-14B-Instruct2026.04 | 39.8 | 34 | 17.1 |