Prompt Injection Defense on AgentDojo Slack suite v1 (test)
61.9CUDelimiting
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DelimitingTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 61.9 | 41.9 | 11.43 | |
| No DefenseTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 57.14 | 46.67 | 12.38 | |
| Repeat PromptTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 52.38 | 38.1 | 5.71 | |
| Tool FilterTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 38.1 | 32.38 | 1.9 | |
| Task ShieldTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 38.1 | 26.67 | 0 | |
| PI DetectorTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 28.57 | 35.24 | 4.76 |