Prompt-Injection Defense on Adaptive Prompt Injection (test)
37Attack Success Rate (ASR)ACT
Evaluation Results
| Method | Links | |
|---|---|---|
| ACTTarget Model=GPT-OSS-20B, Attacker suffixes=best-of-82026.05 | 37 | |
| BCTTarget Model=Phi-4-reasoning, Attacker suffixes=best-of-82026.05 | 42.6 | |
| ACTTarget Model=Phi-4-reasoning, Attacker suffixes=best-of-82026.05 | 49.5 | |
| BCTTarget Model=GPT-OSS-20B, Attacker suffixes=best-of-82026.05 | 57.4 | |
| ACTTarget Model=Qwen3-8B, Attacker suffixes=best-of-82026.05 | 59.3 | |
| ACTTarget Model=Qwen3-1.7B, Attacker suffixes=best-of-82026.05 | 63 | |
| BCTTarget Model=Qwen3-1.7B, Attacker suffixes=best-of-82026.05 | 80.6 | |
| BaselineTarget Model=Phi-4-reasoning, Attacker suffixes=best-of-82026.05 | 90.9 | |
| BCTTarget Model=Qwen3-8B, Attacker suffixes=best-of-82026.05 | 94.4 | |
| BaselineTarget Model=Qwen3-8B, Attacker suffixes=best-of-82026.05 | 96.3 | |
| ACTTarget Model=Gemma-4-E4B-it, Attacker suffixes=best-of-82026.05 | 96.3 | |
| BCTTarget Model=Gemma-4-E4B-it, Attacker suffixes=best-of-82026.05 | 98.1 | |
| BaselineTarget Model=Qwen3-1.7B, Attacker suffixes=best-of-82026.05 | 100 | |
| BaselineTarget Model=GPT-OSS-20B, Attacker suffixes=best-of-82026.05 | 100 | |
| BaselineTarget Model=Gemma-4-E4B-it, Attacker suffixes=best-of-82026.05 | 100 |