Morality Attack (Defense on User Input) on Designed Morality Attacks 1.0 (test)
100RNShieldGemma-9B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| ShieldGemma-9BJustification=false2026.04 | 100 | 88.5 | 100 | 96.8 | 96.3 | — | |
| Llama-Guard-4Justification=false2026.04 | 99.9 | 80 | 99.5 | 69.4 | 87.2 | — | |
| WildGuardJustification=false2026.04 | 96.8 | 77.3 | 94.5 | 79.6 | 87.1 | — | |
| Prompt-Guard-2 (86M)Justification=false2026.04 | 93.2 | 93.5 | 82.3 | 85.3 | 88.6 | — | |
| Granite-Guardian-3.3-8BJustification=true2026.04 | 92.8 | 79.2 | 97.5 | 82.1 | 87.9 | — | |
| Aegis Permissive (CP)Justification=false2026.04 | 65.3 | 31.9 | 74.4 | 12.4 | 46 | — | |
| Aegis Defensive (CP)Justification=false2026.04 | 16 | 5.8 | 41.8 | 2.8 | 16.6 | 69.3 |