Prompt Attack Detection on Curated prompt-attack dataset
0.04LatencyPromptGuard
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| PromptGuardModel Type=Encoder-based, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 0.04 | 40.32 | 15.72 | 22.62 | |
| ProtectAIModel Type=Encoder-based, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 0.04 | 32.14 | 39.62 | 35.49 | |
| gpt_oss_safeguardModel Type=Specialized LLM, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 0.53 | 88.64 | 73.58 | 80.41 | |
| aws_prompt_attackModel Type=Propriertary, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 0.63 | 7.14 | 37.11 | 11.98 | |
| Qwen3Guard (0.6B)Model Type=Specialized LLM, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 1 | 53.75 | 54.09 | 53.92 | |
| gemini-2.5-flash-liteModel Type=LLM Judge, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 1.44 | 81.65 | 81.13 | 81.39 | |
| gemini-2.0-flash-lite-001Model Type=LLM Judge, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 1.52 | 82.14 | 86.79 | 84.4 | |
| gemini-2.5-flashModel Type=LLM Judge, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 1.85 | 77.3 | 89.94 | 83.14 | |
| gemini-3-flash-previewModel Type=LLM Judge, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 2.02 | 79.78 | 91.82 | 85.38 | |
| gpt-5-miniModel Type=LLM Judge, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 2.67 | 89.8 | 83.02 | 86.27 | |
| gpt-5.1Model Type=LLM Judge, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 4.04 | 97.66 | 78.62 | 87.11 | |
| claude-haiku-4-5@20251001Model Type=LLM Judge, Decision threshold (τ)=0.5, Internal reasoning=Disabled2026.03 | 5.88 | 83.53 | 89.31 | 86.32 |