Text-based safety moderation on OpenAI
82.3F1 ScoreGPT-4o
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4oSize=-2025.12 | 82.3 | 88.6 | |
| LLaMA Guard 3Size=8B2025.12 | 81.6 | 85.7 | |
| OMNIGUARD-7BSize=7B2025.12 | 81.1 | 87.9 | |
| Qwen3GuardSize=8B, Mode=loose, Reasoning=false2025.12 | 81 | — | |
| gpt-oss-safeguardSize=20B, Reasoning=true2025.12 | 81 | — | |
| Qwen3-235BSize=235B2025.12 | 80.1 | 85.4 | |
| Qwen2.5-72BSize=72B2025.12 | 79.8 | 85.2 | |
| ShieldGemmaSize=9B, Reasoning=false2025.12 | 79 | — | |
| Llama Guard 3Reasoning=false2025.12 | 79 | — | |
| ThinkGuardSize=8B2025.12 | 78.7 | 79 | |
| OMNIGUARD-3BSize=3B2025.12 | 77.8 | 83.8 | |
| IBM Granite Guardian 3.3Size=8B, Reasoning=false2025.12 | 77 | — | |
| AprielGuardSize=8B, Reasoning=false2025.12 | 77 | — | |
| IBM Granite Guardian 3.3Size=8B, Reasoning=true2025.12 | 77 | — | |
| Llama Guard 2Reasoning=false2025.12 | 76 | — | |
| AprielGuardSize=8B, Reasoning=true2025.12 | 75 | — | |
| LLaMA Guard 2Size=8B2025.12 | 74.4 | 85.6 | |
| Llama Guard 4Reasoning=false2025.12 | 73 | — | |
| IBM Granite Guardian 3.2Size=5B, Reasoning=false2025.12 | 73 | — | |
| Qwen2.5-7BSize=7B2025.12 | 72.6 | 81.4 | |
| Qwen2.5-Omni-7BSize=7B2025.12 | 70 | 70.8 | |
| IBM Granite Guardian 3.1Size=2B, Reasoning=false2025.12 | 69 | — | |
| IBM Granite Guardian 3.2Size=3B, Reasoning=false2025.12 | 68 | — | |
| Qwen3GuardSize=8B, Mode=strict, Reasoning=false2025.12 | 68 | — | |
| LLaMA-3.3-70BSize=70B2025.12 | 58.4 | 80.5 | |
| LLaMA Guard 1Size=7B2025.12 | 32.8 | 74.4 |