Unsafe Prompt Detection on XSTest (test)
87.8PrecisionOpenAI Moderation API
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| OpenAI Moderation API2024.02 | 87.8 | 43 | 57.7 | |
| GPT-4Model version=gpt-4-1106-preview, Zero-shot prompting=true2024.02 | 87.8 | 97 | 92.1 | |
| GradSafe-ZeroBase model=Llama-2-7b-chat-hf, Detection threshold=0.25, Gap threshold=12024.02 | 85.6 | 95 | 90 | |
| Perspective API2024.02 | 83.5 | 33 | 47.3 | |
| Llama GuardBase model=Llama-2 7b, Training=Finetuned on 10,000 prompts2024.02 | 81.3 | 82.5 | 81.9 | |
| Azure API2024.02 | 67.3 | 70 | 68.6 | |
| Llama-2Model version=Llama-2-7b-chat-hf, Zero-shot prompting=true2024.02 | 50.9 | 99 | 67.2 |