Goal Hijacking on Safety-Prompts
92Mean AccuracyScenario-Tailored PC-Inj
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Scenario-Tailored PC-InjLLM=ChatGPT 4o2024.10 | 92 | 0.7 | 13.3 | |
| Scenario-Tailored PC-InjLLM=ChatGPT 4o mini2024.10 | 90.6 | 0.6 | 13.1 | |
| Generalized PC-InjLLM=ChatGPT 4o2024.10 | 89.3 | 0.8 | 10.6 | |
| Generalized PC-InjLLM=ChatGPT 4o mini2024.10 | 88 | 0.5 | 10.5 | |
| Template-Free PC-InjLLM=ChatGPT 4o2024.10 | 85.3 | 0.4 | 6.6 | |
| Scenario-Tailored PC-InjLLM=Qwen 2.52024.10 | 84.5 | 0.4 | 15.2 | |
| Template-Free PC-InjLLM=ChatGPT 4o mini2024.10 | 83.8 | 0.4 | 6.3 | |
| “Ignore the Previous” HijackingLLM=ChatGPT 4o2024.10 | 78.7 | 0.6 | — | |
| Generalized PC-InjLLM=Qwen 2.52024.10 | 78.2 | 0.6 | 8.9 | |
| “Ignore the Previous” HijackingLLM=ChatGPT 4o mini2024.10 | 77.5 | 0.5 | — | |
| Template-Free PC-InjLLM=Qwen 2.52024.10 | 76.5 | 0.7 | 7.2 | |
| “Ignore the Previous” HijackingLLM=Qwen 2.52024.10 | 69.3 | 2.3 | — |