LLM Jailbreaking Attack on AdvBench (100 samples)
100Attack Success RateSemantic Representation Attack
Evaluation Results
| Method | Links | |
|---|---|---|
| Semantic Representation AttackDefense Type=SmoothLLM-Swap2026.05 | 100 | |
| Semantic Representation AttackDefense Type=SmoothLLM-Patch2026.05 | 100 | |
| Semantic Representation AttackDefense Type=SmoothLLM-Insert2026.05 | 96 | |
| Semantic Representation AttackDefense Type=Defenseless2026.05 | 92 | |
| AutoDANDefense Type=Defenseless2026.05 | 90 | |
| AutoDANDefense Type=SmoothLLM-Insert2026.05 | 78 | |
| PAIRDefense Type=Defenseless2026.05 | 76 | |
| AutoDANDefense Type=SmoothLLM-Patch2026.05 | 74 | |
| PAIRDefense Type=SmoothLLM-Insert2026.05 | 62 | |
| AutoDANDefense Type=SmoothLLM-Swap2026.05 | 56 | |
| PAIRDefense Type=SmoothLLM-Patch2026.05 | 52 | |
| PAIRDefense Type=SmoothLLM-Swap2026.05 | 48 |