Harmful Prompt Detection on XSTest
97.44F1 ScoreMLPM
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MLPMBackbone=Llama-8B-Inst, Methodology=Latent-Based2025.02 | 97.44 | — | — | — | — | |
| MLPMBackbone=OLMo2-7B-Inst, Methodology=Latent-Based2025.02 | 96.91 | — | — | — | — | |
| MLPMBackbone=Mistral-7B-Inst, Methodology=Latent-Based2025.02 | 96.1 | — | — | — | — | |
| WildGuardMethodology=Guard Model2025.02 | 95.26 | — | — | — | — | |
| Ayub & MajumdarBackbone=Mistral-7B-Inst, Methodology=Latent-Based2025.02 | 95.04 | — | — | — | — | |
| Abdelnabi et al.Backbone=Mistral-7B-Inst, Methodology=Latent-Based2025.02 | 94.6 | — | — | — | — | |
| Abdelnabi et al.Backbone=OLMo2-7B-Inst, Methodology=Latent-Based2025.02 | 94.43 | — | — | — | — | |
| Ayub & MajumdarBackbone=OLMo2-7B-Inst, Methodology=Latent-Based2025.02 | 94.33 | — | — | — | — | |
| Ayub & MajumdarBackbone=Llama-8B-Inst, Methodology=Latent-Based2025.02 | 92.76 | — | — | — | — | |
| MLPMBackbone=Qwen3-8B-Inst, Methodology=Latent-Based2025.02 | 92.23 | — | — | — | — | |
| Abdelnabi et al.Backbone=Llama-8B-Inst, Methodology=Latent-Based2025.02 | 90.63 | — | — | — | — | |
| Ayub & MajumdarBackbone=Qwen3-8B-Inst, Methodology=Latent-Based2025.02 | 90.21 | — | — | — | — | |
| Llama-3.1-8BDetection Approach=Single-neuron activation, Optimal Thresholding=true2026.05 | 89.6 | 96.9 | 90.2 | 84.8 | 95 | |
| LlamaGuard 3 (8B)Detection Approach=Dedicated 8B classifier2026.05 | 88.8 | 97.5 | 90.2 | 94.9 | 83.4 | |
| LlamaGuard3Methodology=Guard Model2025.02 | 88.52 | — | — | — | — | |
| Abdelnabi et al.Backbone=Qwen3-8B-Inst, Methodology=Latent-Based2025.02 | 88.06 | — | — | — | — | |
| GraniteGuardian-3-1-8BMethodology=Guard Model2025.02 | 85.59 | — | — | — | — | |
| ShieldGemma-9BMethodology=Guard Model2025.02 | 82.41 | — | — | — | — | |
| Aegis-Guard-DMethodology=Guard Model2025.02 | 81.53 | — | — | — | — | |
| Qwen3-32BDetection Approach=Single-neuron activation, Optimal Thresholding=true2026.05 | 80.4 | 90.6 | 83.3 | 84.2 | 77 |