Safety Classification on AdvBench
99.9F1 ScoreGradient-Controlled Decoding (GCD)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Gradient-Controlled Decoding (GCD)2026.04 | 99.9 | 100 | 99.8 | 0 | 0.19 | |
| Safe-Decoding2026.04 | 99.9 | 100 | 99.8 | 0 | 0.19 | |
| GreedyDecoding Strategy=Greedy2026.04 | 99.75 | 100 | 99.5 | 0 | 0.5 | |
| Top-kDecoding Strategy=Top-k2026.04 | 99.75 | 100 | 99.5 | 0 | 0.5 | |
| SCANS2026.04 | 99.62 | 100 | 99.25 | 0 | 0.755 | |
| GradSafe2026.04 | 99.61 | 100 | 99.23 | 0 | 0.77 | |
| Top-pDecoding Strategy=Top-p2026.04 | 99.24 | 100 | 98.5 | 0 | 1.5 | |
| prompt–response verification framework2026.06 | 95.4 | — | — | — | 4.6 | |
| AutoDefense / SelfDefendType=Strongest Applicable Baseline2026.06 | 92.1 | — | — | — | — | |
| Llama 3Model Version=Llama-3.1-8B-Instruct2025.01 | 84 | — | — | — | — | |
| MistralModel Version=Mistral-7B-Instruct-v0.32025.01 | 74 | — | — | — | — | |
| Zephyr RMUModel Version=Zephyr_RMU2025.01 | 70 | — | — | — | — |