Unsafe Prompt Detection on ToxicChat (test)
0.815PrecisionOpenAI Moderation API
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OpenAI Moderation API2024.02 | 0.815 | 0.145 | 0.246 | — | |
| GradSafe-ZeroBase model=Llama-2-7b-chat-hf, Detection threshold=0.25, Gap threshold=12024.02 | 0.753 | 0.667 | 0.707 | — | |
| Llama GuardBase model=Llama-2 7b, Training=Finetuned on 10,000 prompts2024.02 | 0.744 | 0.396 | 0.517 | — | |
| OracleCondition=Perfect Routing2025.02 | 0.65 | 0.881 | 0.748 | 94.17 | |
| Perspective API2024.02 | 0.614 | 0.148 | 0.238 | — | |
| Azure API2024.02 | 0.559 | 0.634 | 0.594 | — | |
| GPT-4Model version=gpt-4-1106-preview, Zero-shot prompting=true2024.02 | 0.475 | 0.831 | 0.604 | — | |
| TSRouting=Thresholding, Small Model=Llama-Guard-3-1B, Large Model=Granite-Guardian-3-8B2025.02 | 0.44 | 0.832 | 0.576 | 83.49 | |
| Granite-Guardian-3-8BModel=Large (Granite-Guardian-3-8B)2025.02 | 0.423 | 0.859 | 0.567 | 93.55 | |
| EntRouting=Entropy-based, Small Model=Llama-Guard-3-1B, Large Model=Granite-Guardian-3-8B2025.02 | 0.418 | 0.716 | 0.528 | 43.52 | |
| SafeRouteRouting=SafeRoute (Ours), Small Model=Llama-Guard-3-1B, Large Model=Granite-Guardian-3-8B2025.02 | 0.395 | 0.781 | 0.525 | 35.12 | |
| CCRouting=Calibration-based, Small Model=Llama-Guard-3-1B, Large Model=Granite-Guardian-3-8B2025.02 | 0.381 | 0.751 | 0.506 | 45.42 | |
| RandomRouting=Random, Small Model=Llama-Guard-3-1B, Large Model=Granite-Guardian-3-8B2025.02 | 0.338 | 0.716 | 0.461 | 52.7 | |
| BCRouting=Bayesian-based, Small Model=Llama-Guard-3-1B, Large Model=Granite-Guardian-3-8B2025.02 | 0.333 | 0.79 | 0.468 | 47.12 | |
| Llama-Guard-3-1BModel=Small (Llama-Guard-3-1B)2025.02 | 0.242 | 0.57 | 0.341 | 17.2 | |
| Llama-2Model version=Llama-2-7b-chat-hf, Zero-shot prompting=true2024.02 | 0.241 | 0.822 | 0.373 | — |