Directional Jailbreak Detection on UltraChat (Safe) × AdvBench (Harmful)
0.971AUROC (mono)Llama-3.1-8B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Llama-3.1-8BProbe Layer=L22, Relative Depth=~69%, Total Layers=32, Direction=downward2026.06 | 0.971 | 0.793 | 0.796 | |
| Qwen3-8BProbe Layer=L25, Relative Depth=~69%, Total Layers=36, Direction=downward2026.06 | 0.939 | 0.826 | 0.838 | |
| Gemma-7bProbe Layer=L19, Relative Depth=~68%, Total Layers=28, Direction=upward2026.06 | 0.741 | 0.808 | 0.813 |