Misuse Detection on Misuse Categories Cybercrime (Phishing)
0.99AUCRepBending
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| RepBendingCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 0.99 | 99 | 1 | |
| GAVELCategory=Classifier, Base Model=Mistral-7B2026.01 | 0.99 | 97 | 0 | |
| Llama Guard 4 (Meta)Category=Moderation, Base Model=Mistral-7B2026.01 | 0.98 | 99 | 0 | |
| Perspective (Google)Category=Moderation, Base Model=Mistral-7B2026.01 | 0.96 | 58 | 0 | |
| Circuit BreakersCategory=Fine-Tuning, Base Model=Mistral-7B2026.01 | 0.89 | 89 | 0 | |
| CASTCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 0.89 | 59 | 80 | |
| Activation ClassifierCategory=Classifier, Base Model=Mistral-7B2026.01 | 0.89 | 82 | 35 | |
| Moderator (OpenAI)Category=Moderation, Base Model=Mistral-7B2026.01 | 0.86 | 86 | 0 | |
| JBShieldCategory=Inference-Time, Base Model=Mistral-7B2026.01 | 0.64 | 52 | 6 |