Jailbreak Attack on SafeBench Tiny
100ASRCOMET
Evaluation Results
| Method | Links | |
|---|---|---|
| COMETVictim Model=GLM-4.5V2026.02 | 100 | |
| COMETVictim Model=Qwen2.5-72B-VL2026.02 | 100 | |
| JailbreakFunctiontarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 100 | |
| SMTtarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 99.67 | |
| Do Anything Nowtarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 98 | |
| Do Anything Nowtarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 98 | |
| TVD-Singletarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 98 | |
| TVD-Singletarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 98 | |
| SMTtarget_llm=Gemini-3-Flash, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 98 | |
| SMTtarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 98 | |
| COMETVictim Model=LlaMa-4-maverick2026.02 | 96 | |
| SMTtarget_llm=Qwen3-Max, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 96 | |
| SMTtarget_llm=DeepSeek-V4-Flash, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 95 | |
| COMETVictim Model=All2026.02 | 94 | |
| JailbreakFunctiontarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 94 | |
| COMETVictim Model=Gemini-2.5-Pro2026.02 | 92 | |
| JailbreakFunctiontarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 92 | |
| HIMRDVictim Model=LlaMa-4-maverick2026.02 | 90 | |
| TVD-Singletarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 90 | |
| SMTtarget_llm=GPT-4o, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 89 | |
| SMTtarget_llm=Avg., prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 88.5 | |
| FigStepVictim Model=Qwen2.5-72B-VL2026.02 | 88 | |
| Crescendotarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 86 | |
| COMETVictim Model=Claude-4.5-Haiku2026.02 | 84 | |
| PAPILLONtarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 82 | |
| Odysseustarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 80 | |
| JailbreakFunctiontarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 80 | |
| ArtPrompttarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 78 | |
| PAPILLONtarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 78 | |
| SMTtarget_llm=Claude-Sonnet-4.5, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 78 | |
| Crescendotarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 76 | |
| TVD-Singletarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 76 | |
| SMTtarget_llm=GPT-5.4, prompting_strategy=one-shot, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 75 | |
| JailbreakFunctiontarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 74.67 | |
| TVD-Singletarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 72.67 | |
| HIMRDVictim Model=GLM-4.5V2026.02 | 72 | |
| ArtPrompttarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 70 | |
| Crescendotarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 70 | |
| Do Anything Nowtarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 68 | |
| CC-BOStarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 68 | |
| CC-BOStarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 66 | |
| FigStepVictim Model=All2026.02 | 65 | |
| FigStepVictim Model=GLM-4.5V2026.02 | 64 | |
| HIMRDVictim Model=Qwen2.5-72B-VL2026.02 | 64 | |
| FigStepVictim Model=LlaMa-4-maverick2026.02 | 62 | |
| PAPILLONtarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 62 | |
| Odysseustarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 62 | |
| Odysseustarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 62 | |
| FigStepVictim Model=Claude-4.5-Haiku2026.02 | 60 | |
| CC-BOStarget_llm=DeepSeek-V4-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 60 | |
| Do Anything Nowtarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 59.67 | |
| CC-BOStarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 59.33 | |
| Crescendotarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 58 | |
| CS-DJVictim Model=GLM-4.5V2026.02 | 56 | |
| HIMRDVictim Model=All2026.02 | 56 | |
| CC-BOStarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 56 | |
| Do Anything Nowtarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 54 | |
| CC-BOStarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 54 | |
| Odysseustarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 54 | |
| PAPILLONtarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 53.67 | |
| FigStepVictim Model=Gemini-2.5-Pro2026.02 | 52 | |
| PAPILLONtarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 52 | |
| CC-BOStarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 52 | |
| Crescendotarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 51 | |
| JailbreakFunctiontarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 50 | |
| Odysseustarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 48.33 | |
| HIMRDVictim Model=Gemini-2.5-Pro2026.02 | 48 | |
| TVD-Singletarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 40 | |
| ArtPrompttarget_llm=Avg., evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 38.67 | |
| TVD-Singletarget_llm=Gemini-3-Flash, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 34 | |
| PAPILLONtarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 32 | |
| JailbreakFunctiontarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 32 | |
| ArtPrompttarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 28 | |
| ArtPrompttarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 28 | |
| Do Anything Nowtarget_llm=GPT-4o, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 26 | |
| CS-DJVictim Model=All2026.02 | 24 | |
| Odysseustarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 24 | |
| CS-DJVictim Model=Gemini-2.5-Pro2026.02 | 22 | |
| CS-DJVictim Model=Qwen2.5-72B-VL2026.02 | 18 | |
| ArtPrompttarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 16 | |
| PAPILLONtarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 16 | |
| Crescendotarget_llm=GPT-5.4, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 16 | |
| CS-DJVictim Model=Claude-4.5-Haiku2026.02 | 14 | |
| Do Anything Nowtarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 14 | |
| ArtPrompttarget_llm=Qwen3-Max, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 12 | |
| CS-DJVictim Model=LlaMa-4-maverick2026.02 | 10 | |
| Odysseustarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 8 | |
| HIMRDVictim Model=Claude-4.5-Haiku2026.02 | 4 | |
| Crescendotarget_llm=Claude-Sonnet-4.5, evaluation_judge=HarmBench-Llama-2-13b-cls2026.07 | 0 |