Guardrail False Positive Rate Estimation on AlpacaEval benign prompts
0False Positive RateLlama Guard
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama GuardBackbone Model=Vicuna-13B-v1.5, Guardrail Processing Stage=Post, Number of NTs (Q)=102026.04 | 0 | |
| Llama GuardBackbone Model=Vicuna-13B-v1.5, Guardrail Processing Stage=Pre, Number of NTs (Q)=102026.04 | 0.37 | |
| SelfGraderBackbone Model=Vicuna-13B-v1.5, Number of NTs (Q)=102026.04 | 0.37 | |
| Perplexity FilterBackbone Model=Vicuna-13B-v1.5, Number of NTs (Q)=102026.04 | 0.62 | |
| Prompt GuardBackbone Model=Vicuna-13B-v1.5, Number of NTs (Q)=102026.04 | 0.62 | |
| SelfDefendBackbone Model=Vicuna-13B-v1.5, Guardrail Processing Stage=Intent, Number of NTs (Q)=102026.04 | 0.86 | |
| SelfDefendBackbone Model=Vicuna-13B-v1.5, Guardrail Processing Stage=Direct, Number of NTs (Q)=102026.04 | 3.6 | |
| GradientCuffBackbone Model=Vicuna-13B-v1.5, Number of NTs (Q)=102026.04 | 13.91 |