Prompt Harmfulness Detection on AegisSafety (test)
99.5F1 ScoreLLaMA Guard 3
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaMA Guard 3Model Size=8B2026.05 | 99.5 | |
| GuardReasonerModel Size=3B2026.05 | 91.39 | |
| GuardReasonerCategory=Reasoning-based Guardrails, Model Size=3B2026.05 | 91.39 | |
| ConsisGuardCategory=Reasoning-based Guardrails, Model Size=7B2026.05 | 91.02 | |
| COLAGUARDModel Size=3B2026.05 | 90.58 | |
| GuardReasonerModel Size=8B2026.05 | 90.18 | |
| GuardReasonerCategory=Reasoning-based Guardrails, Model Size=8B2026.05 | 90.18 | |
| WildGuardCategory=Discriminant-based Guardrails, Model Size=7B2026.05 | 89.69 | |
| COLAGUARDModel Size=8B2026.05 | 89.45 | |
| WildGuardModel Size=7B2026.05 | 89.4 | |
| GuardReasonerModel Size=1B2026.05 | 89.34 | |
| GPT-4o+CoTModel Size=Unknown2026.05 | 88.24 | |
| ConsisGuardCategory=Reasoning-based Guardrails, Model Size=3B2026.05 | 86.58 | |
| GPT-4o+CoTCategory=CoT-Based LLM, Model Size=-2026.05 | 86.32 | |
| Aegis Guard DefensiveModel Size=7B2026.05 | 84.8 | |
| Aegis Guard DefCategory=Discriminant-based Guardrails, Model Size=7B2026.05 | 84.8 | |
| qwen3-CoTCategory=CoT-Based LRM, Model Size=32B2026.05 | 83.36 | |
| o1-previewModel Size=Unknown2026.05 | 83.15 | |
| Aegis Guard PermissiveModel Size=7B2026.05 | 82.9 | |
| Aegis Guard PerCategory=Discriminant-based Guardrails, Model Size=7B2026.05 | 82.9 | |
| Gemini 1.5+CoTCategory=CoT-Based LLM, Model Size=-2026.05 | 82.88 | |
| o1-pre+CoTCategory=CoT-Based LRM, Model Size=-2026.05 | 81.96 | |
| GPT-4oModel Size=Unknown2026.05 | 81.07 | |
| GPT-4+CoTCategory=CoT-Based LLM, Model Size=-2026.05 | 80.52 | |
| QWQ+CoTCategory=CoT-Based LRM, Model Size=32B2026.05 | 80.5 | |
| QwQ-previewModel Size=32B2026.05 | 80.23 | |
| Claude 3.5+CoTCategory=CoT-Based LLM, Model Size=-2026.05 | 78.62 | |
| ShieldGemmaModel Size=9B2026.05 | 77.63 | |
| ShieldGemmaCategory=Discriminant-based Guardrails, Model Size=9B2026.05 | 77.63 | |
| MPNet-based NBFModel Size=115M2025.02 | 74.8 | |
| LLaMA GuardModel Size=7B2025.02 | 74.1 | |
| LLaMA GuardModel Size=7B2026.05 | 74.1 | |
| DistilRoBERTa-based NBFModel Size=87M2025.02 | 74 | |
| LLaMA Guard 2Model Size=8B2026.05 | 71.8 | |
| LLaMA Guard 2Category=Discriminant-based Guardrails, Model Size=8B2026.05 | 71.8 | |
| LLaMA Guard 3Category=Discriminant-based Guardrails, Model Size=8B2026.05 | 71.39 | |
| OpenAI ModerationModel Size=Unknown2025.02 | 31.9 | |
| ModerationCategory=Discriminant-based Guardrails, Model Size=-2026.05 | 31.9 | |
| ShieldGemmaModel Size=2B2025.02 | 7.5 | |
| ShieldGemmaModel Size=2B2026.05 | 7.47 | |
| ShieldGemmaCategory=Discriminant-based Guardrails, Model Size=2B2026.05 | 7.47 |