Safety Classification on OpenAI Moderation
81.4F1 ScoreShieldGemma 27B
Evaluation Results
| Method | Links | |
|---|---|---|
| ShieldGemma 27BOptimization Strategy=Dedicated Safety Guard2026.06 | 81.4 | |
| Label and intent rewardOptimization Strategy=GRPO2026.06 | 80.9 | |
| Label rewardOptimization Strategy=GRPO2026.06 | 79.8 | |
| Gemma-3-12BOptimization Strategy=Zero-shot2026.06 | 79.3 | |
| Human-intent on AIMSOptimization Strategy=Reasoning Distillation, Teacher Model=GPT-OSS-120B, Student Model=Gemma-3-12B2026.06 | 79.2 | |
| GPT-5.4Optimization Strategy=Zero-shot2026.06 | 79.1 | |
| Claude Sonnet 4.6Optimization Strategy=Zero-shot2026.06 | 78.5 | |
| GPT-OSS-Safeguard 120BOptimization Strategy=Dedicated Safety Guard2026.06 | 78 | |
| GPT-OSS-120BOptimization Strategy=Zero-shot2026.06 | 77.5 | |
| IF-DPOOptimization Strategy=DPO2026.06 | 76.6 | |
| LE-DPOOptimization Strategy=DPO2026.06 | 76.5 | |
| Llama-3.1-8BOptimization Strategy=Zero-shot2026.06 | 76.1 | |
| Gemma-3-12BOptimization Strategy=SFT, Supervision Source=Annotated Intents2026.06 | 76.1 | |
| Nemotron Safety 4BOptimization Strategy=Dedicated Safety Guard2026.06 | 74.7 | |
| LlamaGuard 4Optimization Strategy=Dedicated Safety Guard2026.06 | 73.6 | |
| Llama-3.1-8BOptimization Strategy=SFT, Supervision Source=Annotated Intents2026.06 | 72.8 | |
| WildGuard 7BOptimization Strategy=Dedicated Safety Guard2026.06 | 72.4 | |
| GuardReasoner 8BOptimization Strategy=Dedicated Safety Guard2026.06 | 70.4 |