Adversarial Risk Estimation on HarmBench (test)
100ASR@1000Baseline
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| BaselineAttacker=ADV-LLM, Victim=Llama-3.1-8B, Judge=HarmBench Classifier, n=100, N=10002026.01 | 100 | 0 | — | — | — | |
| SABERAttacker=ADV-LLM, Victim=Llama-3.1-8B, Judge=HarmBench Classifier, n=100, N=10002026.01 | 100 | 0 | 0 | — | — | |
| SABERAttacker=ADV-LLM, Victim=Llama-3.1-8B, Judge=LLM Classifier, n=100, N=10002026.01 | 99.81 | 0.5 | 0.82 | — | — | |
| SABERAttacker=Jailbreak-R1, Victim=GPT-4.1-mini, Judge=HarmBench Classifier, n=100, N=10002026.01 | 99.71 | 0.26 | 2.23 | — | — | |
| BaselineAttacker=ADV-LLM, Victim=Llama-3.1-8B, Judge=LLM Classifier, n=100, N=10002026.01 | 97.99 | 1.32 | — | — | — | |
| SABERAttacker=Jailbreak-R1, Victim=GPT-4.1-mini, Judge=LLM Classifier, n=100, N=10002026.01 | 97.93 | 0.5 | 7.62 | — | — | |
| BaselineAttacker=Jailbreak-R1, Victim=GPT-4.1-mini, Judge=HarmBench Classifier, n=100, N=10002026.01 | 97.48 | 2.49 | — | — | — | |
| SABERAttacker=Text Augment., Victim=Llama-3.1-8B, Judge=HarmBench Classifier, n=100, N=10002026.01 | 96.4 | 1.15 | 15.27 | — | — | |
| SABERAttacker=Jailbreak-R1, Victim=Llama-3.1-8B, Judge=HarmBench Classifier, n=100, N=10002026.01 | 96.37 | 2.29 | 9.83 | — | — | |
| SABERAttacker=Jailbreak-R1, Victim=Llama-3.1-8B, Judge=LLM Classifier, n=100, N=10002026.01 | 90.61 | 2.1 | 17.28 | — | — | |
| BaselineAttacker=Jailbreak-R1, Victim=GPT-4.1-mini, Judge=LLM Classifier, n=100, N=10002026.01 | 90.31 | 8.12 | — | — | — | |
| SABERAttacker=Text Augment., Victim=GPT-4.1-mini, Judge=HarmBench Classifier, n=100, N=10002026.01 | 89.44 | 3.18 | 12.21 | — | — | |
| SABERAttacker=Text Augment., Victim=Llama-3.1-8B, Judge=LLM Classifier, n=100, N=10002026.01 | 88.88 | 0.89 | 21.69 | — | — | |
| BaselineAttacker=Jailbreak-R1, Victim=Llama-3.1-8B, Judge=HarmBench Classifier, n=100, N=10002026.01 | 86.54 | 12.12 | — | — | — | |
| SABERAttacker=Text Augment., Victim=GPT-4.1-mini, Judge=LLM Classifier, n=100, N=10002026.01 | 82.54 | 6.05 | 15.5 | — | — | |
| BaselineAttacker=Text Augment., Victim=Llama-3.1-8B, Judge=HarmBench Classifier, n=100, N=10002026.01 | 81.13 | 16.42 | — | — | — | |
| BaselineAttacker=Text Augment., Victim=GPT-4.1-mini, Judge=HarmBench Classifier, n=100, N=10002026.01 | 77.23 | 15.39 | — | — | — | |
| SABERAttacker=ADV-LLM, Victim=GPT-4.1-mini, Judge=HarmBench Classifier, n=100, N=10002026.01 | 74.28 | 0.88 | 10.88 | — | — | |
| BaselineAttacker=Jailbreak-R1, Victim=Llama-3.1-8B, Judge=LLM Classifier, n=100, N=10002026.01 | 73.33 | 19.38 | — | — | — | |
| SABERAttacker=ADV-LLM, Victim=GPT-4.1-mini, Judge=LLM Classifier, n=100, N=10002026.01 | 70.04 | 2.14 | 11.17 | — | — | |
| BaselineAttacker=Text Augment., Victim=GPT-4.1-mini, Judge=LLM Classifier, n=100, N=10002026.01 | 67.04 | 21.55 | — | — | — | |
| BaselineAttacker=Text Augment., Victim=Llama-3.1-8B, Judge=LLM Classifier, n=100, N=10002026.01 | 65.41 | 22.58 | — | — | — | |
| BaselineAttacker=ADV-LLM, Victim=GPT-4.1-mini, Judge=HarmBench Classifier, n=100, N=10002026.01 | 63.4 | 11.76 | — | — | — | |
| BaselineAttacker=ADV-LLM, Victim=GPT-4.1-mini, Judge=LLM Classifier, n=100, N=10002026.01 | 58.87 | 13.31 | — | — | — | |
| BaselineAggregate=Mean across all configs2026.01 | — | — | — | 12.04 | — | |
| SABERAggregate=Mean across all configs2026.01 | — | — | — | 1.66 | 86.2 |