Prompt Injection Defense on AgentDojo Overall v1 (test)
37.11CURepeat Prompt
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Repeat PromptTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 37.11 | 28.3 | 3.82 | |
| No DefenseTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 35.05 | 34.66 | 8.43 | |
| Task ShieldTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 35.05 | 30.05 | 0.95 | |
| DelimitingTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 34.02 | 31.64 | 9.38 | |
| Tool FilterTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 29.9 | 29.57 | 1.43 | |
| PI DetectorTarget Model=GPT-3.5-turbo, Attack=Important Messages2024.12 | 26.8 | 24.8 | 2.86 |