Misuse Detection on Misuse Categories Aggregate Summary
99AUCGAVEL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GAVELCategory=Classifier, Base Model=Mistral-7B2026.01 | 99 | 96 | 0 | |
| Activation ClassifierCategory=Classifier, Base Model=Mistral-7B2026.01 | 97 | 92 | 7 | |
| RepBendingCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 87 | 87 | 2 | |
| Llama Guard 4 (Meta)Category=Moderation, Base Model=Mistral-7B2026.01 | 87 | 93 | 3 | |
| Moderator (OpenAI)Category=Moderation, Base Model=Mistral-7B2026.01 | 69 | 69 | 0 | |
| Circuit BreakersCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 68 | 69 | 6 | |
| CASTCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 68 | 59 | 60 | |
| Perspective (Google)Category=Moderation, Base Model=Mistral-7B2026.01 | 53 | 55 | 2 | |
| JBShieldCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 41 | 63 | 1 |