Safety Guardrail Classification on 630-scenario real-world benchmark (independent set)
95.4Verdict AccuracyAgentTrust v0.5
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| AgentTrust v0.5System Configuration=rules only2026.05 | 95.4 | 2.1 | 6.4 | 2.04 | |
| AgentTrust v0.5 + LLM-JudgeSystem Configuration=hybrid2026.05 | 90.5 | 4.8 | 0.3 | 8.6 | |
| DeepSeek-V3Evaluation Protocol=zero-shot judge2026.05 | 85.1 | 3.2 | 1.7 | 1,271 | |
| NeMo GuardrailsModel Backend=DeepSeek-V32026.05 | 55.1 | 98.4 | 0 | 4,315 | |
| Trivial regex blocklistPatterns=50 patterns2026.05 | 37.9 | 0 | 85.2 | 0.03 |