Morality Attack on Designed Morality Attacks 1.0 (test)
99.8RNShieldGemma-9B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| ShieldGemma-9BJustification=false2026.04 | 99.8 | 83.9 | 99.5 | 94.4 | 94.4 | — | |
| Llama-Guard-4Justification=false2026.04 | 98.8 | 79.2 | 92.9 | 69.5 | 85.1 | — | |
| Claude-Sonnet-4Justification=true2026.04 | 91.6 | 97.2 | 79.2 | 88.9 | 89.2 | — | |
| Gemini-2.5-proJustification=true2026.04 | 91.6 | 87.6 | 97.1 | 88.2 | 91.1 | — | |
| GPT-5Justification=true2026.04 | 90.4 | 96.4 | 71.8 | 88.2 | 86.7 | — | |
| Granite-Guardian-3.3-8BJustification=true2026.04 | 90.1 | 73.3 | 90 | 80.9 | 83.6 | — | |
| Qwen3-235B-A22BJustification=true2026.04 | 85.6 | 94 | 51 | 96.1 | 81.7 | — | |
| WildGuardJustification=false2026.04 | 84.2 | 78.7 | 75.6 | 80.8 | 79.8 | — | |
| DeepSeek-V3.1Justification=true2026.04 | 83.6 | 92 | 56.4 | 96.8 | 82.2 | — | |
| Llama-4-MaverickJustification=true2026.04 | 80.4 | 46.8 | 73.2 | 76.4 | 69.2 | — | |
| GPT-4.1-miniJustification=true2026.04 | 79.2 | 46.4 | 33.5 | 67.5 | 56.7 | — | |
| Agent Permissive(CP)Justification=false2026.04 | 58.8 | 22.4 | 10.6 | 49 | 35.2 | — | |
| Llama-3.1-8BJustification=true2026.04 | 53.2 | 52.8 | 21.4 | 51.8 | 44.8 | — | |
| MDJudgeJustification=true2026.04 | 49.3 | 40.1 | 21.9 | 39.6 | 37.7 | — | |
| Agent Defensive(CP)Justification=false2026.04 | 36.7 | 1.5 | 37.2 | 1.6 | 19.3 | 55.9 |