Misuse Detection on Misuse Categories Psychological Harm (Delusional)
99AUCGAVEL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GAVELCategory=Classifier, Base Model=Mistral-7B2026.01 | 99 | 95 | 0 | |
| Activation ClassifierCategory=Classifier, Base Model=Mistral-7B2026.01 | 98 | 93 | 7 | |
| JBShieldCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 81 | 85 | 3 | |
| Perspective (Google)Category=Moderation, Base Model=Mistral-7B2026.01 | 77 | 50 | 18 | |
| Llama Guard 4 (Meta)Category=Moderation, Base Model=Mistral-7B2026.01 | 62 | 86 | 1 | |
| RepBendingCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 57 | 57 | 1 | |
| Moderator (OpenAI)Category=Moderation, Base Model=Mistral-7B2026.01 | 50 | 50 | 0 | |
| Circuit BreakersCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 49 | 50 | 6 | |
| CASTCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 42 | 47 | 70 |