Misuse Detection on Misuse Categories Scam (Racism)
1AUCGAVEL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GAVELCategory=Classifier, Base Model=Mistral-7B2026.01 | 1 | 0.99 | 0 | |
| CASTCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 0.99 | 0.98 | 0.04 | |
| Moderator (OpenAI)Category=Moderation, Base Model=Mistral-7B2026.01 | 0.99 | 0.99 | 0 | |
| Activation ClassifierCategory=Classifier, Base Model=Mistral-7B2026.01 | 0.98 | 0.95 | 0.02 | |
| RepBendingCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 0.96 | 0.96 | 0.07 | |
| Llama Guard 4 (Meta)Category=Moderation, Base Model=Mistral-7B2026.01 | 0.95 | 0.94 | 0.07 | |
| Perspective (Google)Category=Moderation, Base Model=Mistral-7B2026.01 | 0.89 | 0.62 | 0.01 | |
| Circuit BreakersCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 0.87 | 0.88 | 0.23 | |
| JBShieldCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 0.69 | 0.81 | 0 |