Agent Action Safety Verification on internal benchmark 300-scenario
95Verdict AccuracyAgentTrust v0.5
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| AgentTrust v0.5Evaluation Protocol=rules only2026.05 | 95 | 2.3 | 5.4 | 1.72 | |
| AgentTrust v0.5 + LLM-JudgeEvaluation Protocol=hybrid2026.05 | 88 | 9 | 0.8 | 1,461 | |
| DeepSeek-V3Evaluation Protocol=zero-shot judge2026.05 | 82.3 | 7.5 | 2.3 | 1,345 | |
| Trivial regex blocklistPattern Count=50 patterns2026.05 | 49.3 | 0 | 88.4 | 0.05 | |
| NeMo GuardrailsBackend LLM=DeepSeek-V32026.05 | 44.7 | 96.2 | 0 | 3,558 |