Guardrail False Positive Rate Estimation on OR-Bench benign prompts
0False Positive RatePerplexity Filter
Evaluation Results
| Method | Links | |
|---|---|---|
| Perplexity FilterBackbone Model=Vicuna-13B-v1.5, Number of NTs (Q)=102026.04 | 0 | |
| Llama GuardBackbone Model=Vicuna-13B-v1.5, Guardrail Processing Stage=Post, Number of NTs (Q)=102026.04 | 1 | |
| SelfGraderBackbone Model=Vicuna-13B-v1.5, Number of NTs (Q)=102026.04 | 1.6 | |
| Prompt GuardBackbone Model=Vicuna-13B-v1.5, Number of NTs (Q)=102026.04 | 2.5 | |
| Llama GuardBackbone Model=Vicuna-13B-v1.5, Guardrail Processing Stage=Pre, Number of NTs (Q)=102026.04 | 5.4 | |
| SelfDefendBackbone Model=Vicuna-13B-v1.5, Guardrail Processing Stage=Intent, Number of NTs (Q)=102026.04 | 8.6 | |
| GradientCuffBackbone Model=Vicuna-13B-v1.5, Number of NTs (Q)=102026.04 | 12.4 | |
| SelfDefendBackbone Model=Vicuna-13B-v1.5, Guardrail Processing Stage=Direct, Number of NTs (Q)=102026.04 | 21.6 |