Detection-based Defense on AgentDojo
0False Positive Rate (FPR)AgentShield
Evaluation Results
| Method | Links | |
|---|---|---|
| AgentShieldLanguages=EN,KU,AR,CS, Overhead=<1%, Extra LLM=No, Labels=No2026.05 | 0 | |
| PromptArmorLanguages=EN, Overhead=+1 LLM/tool, Extra LLM=Yes, Labels=No2026.05 | 1 | |
| MELONLanguages=EN, Overhead=2× inference, Extra LLM=Yes, Labels=No2026.05 | 9.28 |