Unsafe Prompt Detection on ToxicChat
75.5AUPRCGradSafe-Zero
Evaluation Results
| Method | Links | |
|---|---|---|
| GradSafe-ZeroBackbone=Llama-2-7b-chat-hf2024.02 | 75.5 | |
| Llama GuardBackbone=Llama-2 7b2024.02 | 63.5 | |
| OpenAI Moderation API2024.02 | 60.4 | |
| Perspective API2024.02 | 48.7 |
| Method | Links | |
|---|---|---|
| GradSafe-ZeroBackbone=Llama-2-7b-chat-hf2024.02 | 75.5 | |
| Llama GuardBackbone=Llama-2 7b2024.02 | 63.5 | |
| OpenAI Moderation API2024.02 | 60.4 | |
| Perspective API2024.02 | 48.7 |