Content Safety Detection on HarmBench
95.6F1 Scoreprompt–response verification framework
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| prompt–response verification framework2026.06 | 95.6 | 3.9 | |
| WildGuard / ShieldGemmaType=Strongest Applicable Baseline2026.06 | 90.3 | — |
| Method | Links | ||
|---|---|---|---|
| prompt–response verification framework2026.06 | 95.6 | 3.9 | |
| WildGuard / ShieldGemmaType=Strongest Applicable Baseline2026.06 | 90.3 | — |