Misuse Detection on Misuse Categories Scam (Elections)
0.99AUCGAVEL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GAVELCategory=Classifier, Base Model=Mistral-7B2026.01 | 0.99 | 98 | 2 | |
| Activation ClassifierCategory=Classifier, Base Model=Mistral-7B2026.01 | 0.9 | 75 | 12 | |
| Perspective (Google)Category=Moderation, Base Model=Mistral-7B2026.01 | 0.89 | 50 | 1 | |
| CASTCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 0.82 | 66 | 67 | |
| Llama Guard 4 (Meta)Category=Moderation, Base Model=Mistral-7B2026.01 | 0.79 | 88 | 7 | |
| RepBendingCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 0.5 | 50 | 0 | |
| Moderator (OpenAI)Category=Moderation, Base Model=Mistral-7B2026.01 | 0.5 | 50 | 0 | |
| Circuit BreakersCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 0.42 | 42 | 15 | |
| JBShieldCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 0.39 | 53 | 0 |