Misuse Detection on Misuse Categories Psychological Harm (Anti-LGBTQ)
100AUCModerator (OpenAI)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Moderator (OpenAI)Category=Moderation, Base Model=Mistral-7B2026.01 | 100 | 100 | 0 | |
| GAVELCategory=Classifier, Base Model=Mistral-7B2026.01 | 100 | 100 | 0 | |
| RepBendingCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 99 | 99 | 1 | |
| CASTCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 99 | 91 | 17 | |
| Llama Guard 4 (Meta)Category=Moderation, Base Model=Mistral-7B2026.01 | 99 | 99 | 1 | |
| Activation ClassifierCategory=Classifier, Base Model=Mistral-7B2026.01 | 99 | 98 | 3 | |
| Circuit BreakersCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 94 | 95 | 9 | |
| JBShieldCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 73 | 84 | 1 | |
| Perspective (Google)Category=Moderation, Base Model=Mistral-7B2026.01 | 8 | 62 | 0 |