Adversarial Attack Detection on Agentic AI workflow benchmark internal
100PrecisionQwen3Guard-8B-loose
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen3Guard-8B-looseReasoning=Without2025.12 | 100 | 1 | 1 | 0 | |
| AprielGuard-8BReasoning=Without2025.12 | 100 | 91 | 95 | 1 | |
| AprielGuard-8BReasoning=With2025.12 | 99 | 84 | 91 | 2 | |
| Qwen3Guard-8B-strictReasoning=Without2025.12 | 98 | 13 | 22 | 1 | |
| gpt-oss-safeguard-20BReasoning=With2025.12 | 98 | 19 | 32 | 1 | |
| Llama-guard-4-12BReasoning=Without2025.12 | 91 | 20 | 33 | 27 | |
| IBM-Granite-Guardian-3.3-8BReasoning=Without2025.12 | 91 | 21 | 34 | 6 | |
| IBM-Granite-Guardian-3.3-8BReasoning=With2025.12 | 88 | 17 | 28 | 6 | |
| Llama-guard-3-8BReasoning=Without2025.12 | 86 | 18 | 30 | 8 | |
| IBM-Granite-Guardian-3.1-2BReasoning=Without2025.12 | 83 | 11 | 19 | 6 | |
| IBM-Granite-Guardian-3.2-3BReasoning=Without2025.12 | 78 | 81 | 80 | 62 | |
| Llama-Prompt-Guard-2-86MReasoning=Without2025.12 | 74 | 75 | 74 | 71 | |
| ShieldGemma-9BReasoning=Without2025.12 | 72 | 1 | 3 | 2 | |
| IBM-Granite-Guardian-3.2-5BReasoning=Without2025.12 | 72 | 82 | 77 | 90 |