Attack Success Rate (ASR) on JailbreakBench
0Attack Success Rate (ASR)Crescendo
Evaluation Results
| Method | Links | |
|---|---|---|
| Crescendotarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 0 | |
| Odysseustarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 0 | |
| ArtPrompttarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 1 | |
| PAPILLONtarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 2 | |
| JailbreakFunctiontarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 4 | |
| GCGModel=Llama-3.1 8B, Config=100 steps, Time=15 min2026.06 | 5 | |
| ArtPrompttarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 9 | |
| Do Anything Nowtarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 9 | |
| PAPILLONtarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 9 | |
| ArtPrompttarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 11 | |
| Crescendotarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 13 | |
| Odysseustarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 16 | |
| Do Anything Nowtarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 23 | |
| TVD-Singletarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 28 | |
| PAPILLONtarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 29 | |
| CC-BOStarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 30 | |
| ArtPrompttarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 31.17 | |
| PAPILLONtarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 31.33 | |
| Odysseustarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 34 | |
| Crescendotarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 35.5 | |
| TVD-Singletarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 38 | |
| Odysseustarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 38 | |
| Odysseustarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 39 | |
| PAPILLONtarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 40 | |
| Do Anything Nowtarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 42 | |
| CC-BOStarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 42 | |
| Crescendotarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 43 | |
| CC-BOStarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 44.17 | |
| GCGModel=Qwen-2.5 7B, Config=100 steps, Time=15 min2026.06 | 45 | |
| CC-BOStarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 45 | |
| Crescendotarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 46 | |
| CC-BOStarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 47 | |
| JailbreakFunctiontarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 48 | |
| ArtPrompttarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 49 | |
| CC-BOStarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 49 | |
| PAPILLONtarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 51 | |
| CC-BOStarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 52 | |
| Do Anything Nowtarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 52.83 | |
| Crescendotarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 53 | |
| ArtPrompttarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 54 | |
| PAPILLONtarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 57 | |
| Crescendotarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 58 | |
| Do Anything Nowtarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 61 | |
| ArtPrompttarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 63 | |
| Odysseustarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 64 | |
| TVD-Singletarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 70.83 | |
| JailbreakFunctiontarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 70.83 | |
| TVD-Singletarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 73 | |
| CLSModel=Qwen-2.5 7B, Config=α = 2, Time=1 sec2026.06 | 75 | |
| SMTtarget_llm=GPT-5.4, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 77 | |
| Do Anything Nowtarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 82 | |
| Odysseustarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 82 | |
| SMTtarget_llm=Claude-Sonnet-4.5, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 83 | |
| CLSModel=Llama-3.1 8B, Config=α = 2, Time=1 sec2026.06 | 85 | |
| JailbreakFunctiontarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 86 | |
| TVD-Singletarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 87 | |
| CLSModel=Qwen-2.5 7B, Config=α = 3, Time=1 sec2026.06 | 90 | |
| SMTtarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 92 | |
| SMTtarget_llm=Avg., prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 92.83 | |
| JailbreakFunctiontarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 93 | |
| CLSModel=Llama-3.1 8B, Config=α = 3, Time=1 sec2026.06 | 95 | |
| JailbreakFunctiontarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 96 | |
| JailbreakFunctiontarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 98 | |
| SMTtarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 98 | |
| SMTtarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 98.33 | |
| TVD-Singletarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 99 | |
| SMTtarget_llm=GPT-4o, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 99 | |
| SMTtarget_llm=Qwen3-Max, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 99 | |
| SMTtarget_llm=DeepSeek-V4-Flash, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 99 | |
| Do Anything Nowtarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| TVD-Singletarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=Gemini-3-Flash, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 |