Harmful prompt detection on OpenAI
81.3F1 ScoreStage1-SFT-v2
Evaluation Results
| Method | Links | |
|---|---|---|
| Stage1-SFT-v2Params=4B2026.07 | 81.3 | |
| Stage1-SFT-v4Params=4B2026.07 | 80.3 | |
| Stage2-SFTParams=4B2026.07 | 79.8 | |
| LlamaGuard3Methodology=Guard Model2025.02 | 79.11 | |
| Stage3-DPOParams=4B2026.07 | 79 | |
| GraniteGuardian-3-1-8BMethodology=Guard Model2025.02 | 77.63 | |
| ShieldGemma-9BMethodology=Guard Model2025.02 | 77.63 | |
| Stage1-SFT-v1Params=4B2026.07 | 76.8 | |
| Aegis-Guard-DMethodology=Guard Model2025.02 | 76.44 | |
| Stage1-SFT-v3Params=4B2026.07 | 75.8 | |
| Ayub & MajumdarBackbone=Mistral-7B-Inst, Methodology=Latent-Based2025.02 | 75.69 | |
| Abdelnabi et al.Backbone=OLMo2-7B-Inst, Methodology=Latent-Based2025.02 | 75.19 | |
| YuFeng-XGuard-Reason-8BParams=8B2026.07 | 74.7 | |
| Ayub & MajumdarBackbone=Qwen3-8B-Inst, Methodology=Latent-Based2025.02 | 74.56 | |
| MLPMBackbone=OLMo2-7B-Inst, Methodology=Latent-Based2025.02 | 74.21 | |
| Stage1-SFT-v0Params=4B2026.07 | 74 | |
| MLPMBackbone=Llama-8B-Inst, Methodology=Latent-Based2025.02 | 72.85 | |
| YuFeng-XGuard-Reason-0.6BParams=0.6B2026.07 | 72.8 | |
| MLPMBackbone=Qwen3-8B-Inst, Methodology=Latent-Based2025.02 | 72.35 | |
| WildGuardMethodology=Guard Model2025.02 | 72.28 | |
| MLPMBackbone=Mistral-7B-Inst, Methodology=Latent-Based2025.02 | 70.68 | |
| Qwen3Guard-8B-GenParams=8B2026.07 | 68.5 | |
| Abdelnabi et al.Backbone=Qwen3-8B-Inst, Methodology=Latent-Based2025.02 | 68.45 | |
| Qwen3Guard-4B-GenParams=4B2026.07 | 68.4 | |
| Abdelnabi et al.Backbone=Llama-8B-Inst, Methodology=Latent-Based2025.02 | 67.99 | |
| Ayub & MajumdarBackbone=OLMo2-7B-Inst, Methodology=Latent-Based2025.02 | 66.95 | |
| Ayub & MajumdarBackbone=Llama-8B-Inst, Methodology=Latent-Based2025.02 | 66.6 | |
| Qwen3Guard-0.8B-GenParams=0.8B2026.07 | 66.2 | |
| Abdelnabi et al.Backbone=Mistral-7B-Inst, Methodology=Latent-Based2025.02 | 64.59 |