Red-Teaming (ASR) on DANGEROUSQA
0ASRGPT-4
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GPT-4Prompting Strategy=STANDARD2023.08 | 0 | — | — | — | — | |
| GPT-4Prompting Strategy=COT2023.08 | 0 | — | — | — | — | |
| CHATGPTPrompting Strategy=STANDARD2023.08 | 0 | — | — | — | — | |
| SafeTransformerPrompt=Standard, Inference Mode=eos2026.03 | 0 | — | — | — | — | |
| SafeTransformerPrompt=CoT, Inference Mode=eos2026.03 | 0 | — | — | — | — | |
| SafeTransformerPrompt=CoU, Inference Mode=eos2026.03 | 0 | — | — | — | — | |
| SafeTransformerPrompt=Suffix, Inference Mode=eos2026.03 | 0 | — | — | — | — | |
| SafeTransformerPrompt=Avg., Inference Mode=eos2026.03 | 0 | — | — | — | — | |
| SafeTransformerPrompt=Standard, Inference Mode=avg2026.03 | 0 | — | — | — | — | |
| SafeTransformerPrompt=CoT, Inference Mode=avg2026.03 | 0 | — | — | — | — | |
| SafeTransformerPrompt=CoU, Inference Mode=avg2026.03 | 0 | — | — | — | — | |
| SafeTransformerPrompt=Suffix, Inference Mode=avg2026.03 | 0 | — | — | — | — | |
| SafeTransformerPrompt=Avg., Inference Mode=avg2026.03 | 0 | — | — | — | — | |
| CHATGPTPrompting Strategy=COT2023.08 | 0.5 | — | — | — | — | |
| STARLING (BLUE)Prompting Strategy=STANDARD2023.08 | 1.5 | — | — | — | — | |
| VICUNA-7BPrompting Strategy=STANDARD2023.08 | 2.5 | — | — | — | — | |
| STABLEBELUGA-13BPrompting Strategy=STANDARD2023.08 | 2.6 | — | — | — | — | |
| VICUNA-13BPrompting Strategy=STANDARD2023.08 | 2.7 | — | — | — | — | |
| Llama-3.2-1B (SFT)Prompt=Suffix2026.03 | 3.5 | — | — | — | — | |
| STARLING (BLUE-RED)Prompting Strategy=STANDARD2023.08 | 5 | — | — | — | — | |
| Llama-3.2-1BPrompt=Suffix2026.03 | 5 | — | — | — | — | |
| Llama-3.2-1B (SFT)Prompt=Standard2026.03 | 6.5 | — | — | — | — | |
| VICUNA-FT-7BPrompting Strategy=STANDARD2023.08 | 9.5 | — | — | — | — | |
| STABLEBELUGA-7BPrompting Strategy=STANDARD2023.08 | 10.2 | — | — | — | — | |
| Llama-3.2-1BPrompt=Standard2026.03 | 13 | — | — | — | — | |
| Llama-3.2-1B (SFT)Prompt=Avg.2026.03 | 14.25 | — | — | — | — | |
| Llama-3.2-1BPrompt=CoU2026.03 | 16.5 | — | — | — | — | |
| Llama-3.2-1BPrompt=Avg.2026.03 | 17 | — | — | — | — | |
| Llama-3.2-1B (SFT)Prompt=CoT2026.03 | 19 | — | — | — | — | |
| Llama-3.2-1B (SFT)Prompt=CoU2026.03 | 28 | — | — | — | — | |
| Llama-3.2-1BPrompt=CoT2026.03 | 33.5 | — | — | — | — | |
| GPT-4Prompting Strategy=RED-EVAL2023.08 | 36.7 | — | — | — | — | |
| VICUNA-FT-7BPrompting Strategy=COT2023.08 | 46.5 | — | — | — | — | |
| STARLING (BLUE)Prompting Strategy=COT2023.08 | 48.5 | — | — | — | — | |
| VICUNA-13BPrompting Strategy=COT2023.08 | 49 | — | — | — | — | |
| VICUNA-7BPrompting Strategy=COT2023.08 | 53.2 | — | — | — | — | |
| STARLING (BLUE-RED)Prompting Strategy=COT2023.08 | 57 | — | — | — | — | |
| STABLEBELUGA-13BPrompting Strategy=COT2023.08 | 63 | — | — | — | — | |
| LLAMA2-FT-7BPrompting Strategy=STANDARD2023.08 | 72.2 | — | — | — | — | |
| CHATGPTPrompting Strategy=RED-EVAL2023.08 | 73.6 | — | — | — | — | |
| STABLEBELUGA-7BPrompting Strategy=COT2023.08 | 75.5 | — | — | — | — | |
| STABLEBELUGA-13BPrompting Strategy=RED-EVAL2023.08 | 81.5 | — | — | — | — | |
| STARLING (BLUE)Prompting Strategy=RED-EVAL2023.08 | 82.5 | — | — | — | — | |
| VICUNA-FT-7BPrompting Strategy=RED-EVAL2023.08 | 83.5 | — | — | — | — | |
| LLAMA2-FT-7BPrompting Strategy=COT2023.08 | 86 | — | — | — | — | |
| STARLING (BLUE-RED)Prompting Strategy=RED-EVAL2023.08 | 86.5 | — | — | — | — | |
| VICUNA-13BPrompting Strategy=RED-EVAL2023.08 | 87 | — | — | — | — | |
| LLAMA2-FT-7BPrompting Strategy=RED-EVAL2023.08 | 90 | — | — | — | — | |
| STABLEBELUGA-7BPrompting Strategy=RED-EVAL2023.08 | 90.5 | — | — | — | — | |
| VICUNA-7BPrompting Strategy=RED-EVAL2023.08 | 91.5 | — | — | — | — | |
| CHATGPTPrompting Strategy=Overall2023.08 | — | 24.7 | — | — | — | |
| CHATGPT2023.08 | — | 24.4 | 0 | 0.005 | 72.8 | |
| GPT-4Prompting Strategy=Overall2023.08 | — | 12.2 | — | — | — | |
| GPT-42023.08 | — | 21.7 | — | — | 65.1 | |
| LLAMA2-FTParameters=7B, fine-tuned=true2023.08 | — | 82.6 | 0.722 | 0.86 | 89.6 | |
| LLAMA2-FT-7BPrompting Strategy=Overall2023.08 | — | 82.7 | — | — | — | |
| STABLEBELUGAParameters=13B2023.08 | — | 52.3 | 0.026 | 0.63 | 91.5 | |
| STABLEBELUGAParameters=7B2023.08 | — | 59 | 0.102 | 0.755 | 91.5 | |
| STABLEBELUGA-13BPrompting Strategy=Overall2023.08 | — | 49 | — | — | — | |
| STABLEBELUGA-7BPrompting Strategy=Overall2023.08 | — | 58.7 | — | — | — | |
| STARLINGData=BLUE, Strategy=A2023.08 | — | 42.1 | 0.015 | 0.485 | 76.5 | |
| STARLINGData=BLUE-RED, Strategy=B2023.08 | — | 49.2 | 0.05 | 0.57 | 85.5 | |
| STARLING (BLUE-RED)Prompting Strategy=Overall2023.08 | — | 49.5 | — | — | — | |
| STARLING (BLUE)Prompting Strategy=Overall2023.08 | — | 44.1 | — | — | — | |
| VICUNAParameters=13B2023.08 | — | 45 | 0.027 | 0.49 | 83.5 | |
| VICUNAParameters=7B2023.08 | — | 47.7 | 0.025 | 0.532 | 87.5 | |
| VICUNA-13BPrompting Strategy=Overall2023.08 | — | 46.2 | — | — | — | |
| VICUNA-7BPrompting Strategy=Overall2023.08 | — | 49 | — | — | — | |
| VICUNA-FTParameters=7B, fine-tuned=true2023.08 | — | 47.3 | 0.095 | 0.465 | 86 | |
| VICUNA-FT-7BPrompting Strategy=Overall2023.08 | — | 46.5 | — | — | — |