LLM Inference Efficiency on HarmBench-Standard and AdvBench
2SlowdownALIGNBEAM
Evaluation Results
| Method | Links | |
|---|---|---|
| ALIGNBEAMK budget=1, Draft models=Llama-3.1-8B / Qwen3-8B-Base, Anchor model=Qwen2.5-3B-Instruct2026.06 | 2 | |
| Hard prefixDraft models=Llama-3.1-8B / Qwen3-8B-Base, Anchor model=Qwen2.5-3B-Instruct2026.06 | 2.4 | |
| RAINDraft models=Llama-3.1-8B / Qwen3-8B-Base, Anchor model=Qwen2.5-3B-Instruct2026.06 | 3.78 | |
| LlamaGuard-responseK=3, Draft models=Llama-3.1-8B / Qwen3-8B-Base, Anchor model=Qwen2.5-3B-Instruct2026.06 | 3.9 | |
| LlamaGuard-promptK=3, Draft models=Llama-3.1-8B / Qwen3-8B-Base, Anchor model=Qwen2.5-3B-Instruct2026.06 | 4.1 | |
| ALIGNBEAMK=3 default, Draft models=Llama-3.1-8B / Qwen3-8B-Base, Anchor model=Qwen2.5-3B-Instruct2026.06 | 4.6 | |
| Top-k contrastive decodingBase model=Qwen3-8B-Base, Draft models=Llama-3.1-8B / Qwen3-8B-Base, Anchor model=Qwen2.5-3B-Instruct2026.06 | 4.8 | |
| Proxy TuningBase model=Qwen3-8B-Base, Draft models=Llama-3.1-8B / Qwen3-8B-Base, Anchor model=Qwen2.5-3B-Instruct2026.06 | 4.9 |