Harmful prompt detection on HarmB
100F1 ScoreMLPM
Evaluation Results
| Method | Links | |
|---|---|---|
| MLPMBackbone=Qwen3-8B-Inst, Methodology=Latent-Based2025.02 | 100 | |
| Qwen3Guard-GenSize=4B2026.07 | 100 | |
| Qwen3Guard-GenSize=8B2026.07 | 100 | |
| MLPMBackbone=Llama-8B-Inst, Methodology=Latent-Based2025.02 | 99.58 | |
| Abdelnabi et al.Backbone=Qwen3-8B-Inst, Methodology=Latent-Based2025.02 | 99.37 | |
| WildGuardMethodology=Guard Model2025.02 | 99.37 | |
| HaloGuard 1.0Size=4B2026.07 | 99.2 | |
| MLPMBackbone=Mistral-7B-Inst, Methodology=Latent-Based2025.02 | 99.16 | |
| LlamaGuard3Methodology=Guard Model2025.02 | 98.94 | |
| WildGuardSize=7B2026.07 | 98.9 | |
| PolyGuard-QwenSize=7B2026.07 | 98.7 | |
| Qwen3Guard-GenSize=0.6B2026.07 | 98.7 | |
| HaloGuard 1.0Size=0.8B2026.07 | 98.7 | |
| MLPMBackbone=OLMo2-7B-Inst, Methodology=Latent-Based2025.02 | 98.51 | |
| Abdelnabi et al.Backbone=Mistral-7B-Inst, Methodology=Latent-Based2025.02 | 97.86 | |
| LlamaGuard4Size=12B2026.07 | 97.2 | |
| Ayub & MajumdarBackbone=Llama-8B-Inst, Methodology=Latent-Based2025.02 | 96.98 | |
| Ayub & MajumdarBackbone=OLMo2-7B-Inst, Methodology=Latent-Based2025.02 | 96.54 | |
| Abdelnabi et al.Backbone=Llama-8B-Inst, Methodology=Latent-Based2025.02 | 95.18 | |
| Abdelnabi et al.Backbone=OLMo2-7B-Inst, Methodology=Latent-Based2025.02 | 93.3 | |
| Ayub & MajumdarBackbone=Mistral-7B-Inst, Methodology=Latent-Based2025.02 | 90.87 | |
| Ayub & MajumdarBackbone=Qwen3-8B-Inst, Methodology=Latent-Based2025.02 | 90.62 | |
| GraniteGuardian-3-1-8BMethodology=Guard Model2025.02 | 79.9 | |
| NemoGuardSize=8B2026.07 | 75.2 | |
| Aegis-Guard-DMethodology=Guard Model2025.02 | 70.46 | |
| ShieldGemma-9BMethodology=Guard Model2025.02 | 69.04 | |
| ShieldGemmaSize=27B2026.07 | 57.3 |