Directional Jailbreak Detection on UltraChat (Safe) × HarmBench (Harmful)
0.956AUROC (mono)Llama-3.1-8B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Llama-3.1-8BProbe Layer=L22, Relative Depth=~69%, Total Layers=32, Direction=downward2026.06 | 0.956 | 0.773 | 0.777 | |
| Qwen3-8BProbe Layer=L25, Relative Depth=~69%, Total Layers=36, Direction=downward2026.06 | 0.928 | 0.76 | 0.772 | |
| Gemma-7bProbe Layer=L19, Relative Depth=~68%, Total Layers=28, Direction=upward2026.06 | 0.725 | 0.798 | 0.797 |