Harmfulness Detection on OpenAI Moderation
92.9Macro F1 ScoreSIREN
Evaluation Results
| Method | Links | |
|---|---|---|
| SIRENBackbone=Llama3.2-1B2026.04 | 92.9 | |
| SIRENBackbone=Llama3.1-8B2026.04 | 92 | |
| SIRENBackbone=Qwen3-0.6B2026.04 | 91.3 | |
| SIRENBackbone=Qwen3-4B2026.04 | 91.2 | |
| LlamaGuard3Backbone=Llama3.1-8B2026.04 | 85.3 | |
| Qwen3-VL-235B2026.06 | 82.19 | |
| ShieldGemma-9B2026.06 | 81.43 | |
| Aegis Guard 2.0Model Size=8B2026.05 | 81 | |
| GPT-5.12026.06 | 80.24 | |
| LLaMA Guard 3Model Size=8B2026.05 | 79.69 | |
| LLaMA Guard 3Category=Discriminant-based Guardrails, Model Size=8B2026.05 | 79.69 | |
| Qwen3Guard-8BPolicy=loose2026.06 | 79.58 | |
| Llama Guard 32026.06 | 79.04 | |
| ModerationCategory=Discriminant-based Guardrails, Model Size=-2026.05 | 79 | |
| ShieldGemmaModel Size=9B2026.05 | 78.58 | |
| ShieldGemmaCategory=Discriminant-based Guardrails, Model Size=9B2026.05 | 78.58 | |
| Qwen3GuardBackbone=Qwen3-4B2026.04 | 78.3 | |
| GPT-4o+CoTModel Size=Unknown2026.05 | 76.78 | |
| GPT-4o+CoTCategory=CoT-Based LLM, Model Size=-2026.05 | 76.2 | |
| LLaMA Guard 2Model Size=8B2026.05 | 76.1 | |
| LLaMA Guard 2Category=Discriminant-based Guardrails, Model Size=8B2026.05 | 76.1 | |
| Qwen3GuardBackbone=Qwen3-0.6B2026.04 | 75.9 | |
| LLaMA GuardModel Size=7B2026.05 | 75.8 | |
| SingGuard-8BParameters=8B2026.06 | 75.61 | |
| Gemini3-Pro2026.06 | 75.45 | |
| o1-pre+CoTCategory=CoT-Based LRM, Model Size=-2026.05 | 75.24 | |
| Aegis Guard PermissiveModel Size=7B2026.05 | 74.7 | |
| Aegis Guard PerCategory=Discriminant-based Guardrails, Model Size=7B2026.05 | 74.7 | |
| o1-previewModel Size=Unknown2026.05 | 74.6 | |
| ConsisGuardCategory=Reasoning-based Guardrails, Model Size=3B2026.05 | 74.35 | |
| SingGuard-4BParameters=4B2026.06 | 73.6 | |
| COLAGUARDModel Size=8B2026.05 | 73.45 | |
| COLAGUARDModel Size=3B2026.05 | 73.15 | |
| ConsisGuardCategory=Reasoning-based Guardrails, Model Size=7B2026.05 | 73.11 | |
| qwen3-CoTCategory=CoT-Based LRM, Model Size=32B2026.05 | 72.89 | |
| WildGuard2026.06 | 72.52 | |
| WildGuardModel Size=7B2026.05 | 72.1 | |
| WildGuardCategory=Discriminant-based Guardrails, Model Size=7B2026.05 | 72.1 | |
| GuardReasonerModel Size=8B2026.05 | 72 | |
| GuardReasonerCategory=Reasoning-based Guardrails, Model Size=8B2026.05 | 72 | |
| GuardReasonerModel Size=3B2026.05 | 71.87 | |
| GuardReasonerCategory=Reasoning-based Guardrails, Model Size=3B2026.05 | 71.87 | |
| YuFeng-XGuard-Reason-8B2026.06 | 71.87 | |
| SingGuard-2BParameters=2B2026.06 | 71.59 | |
| GuardReasoner-VL-7B2026.06 | 71.24 | |
| GuardReasonerModel Size=1B2026.05 | 70.06 | |
| Qwen3Guard-8BPolicy=strict2026.06 | 68.04 | |
| LlamaGuard3Backbone=Llama3.2-1B2026.04 | 67.5 | |
| Aegis Guard DefensiveModel Size=7B2026.05 | 67.5 | |
| Aegis Guard DefCategory=Discriminant-based Guardrails, Model Size=7B2026.05 | 67.5 | |
| QWQ+CoTCategory=CoT-Based LRM, Model Size=32B2026.05 | 65.34 | |
| GPT-4+CoTCategory=CoT-Based LLM, Model Size=-2026.05 | 63.65 | |
| GPT-4oModel Size=Unknown2026.05 | 62.26 | |
| Gemini 1.5+CoTCategory=CoT-Based LLM, Model Size=-2026.05 | 62.06 | |
| QwQ-previewModel Size=32B2026.05 | 61.58 | |
| GraniteGuardian2026.06 | 56.84 | |
| Claude 3.5+CoTCategory=CoT-Based LLM, Model Size=-2026.05 | 53.72 | |
| ShieldGemmaModel Size=2B2026.05 | 13.89 | |
| ShieldGemmaCategory=Discriminant-based Guardrails, Model Size=2B2026.05 | 13.89 |