Jailbreak Robustness on WildJailbreak
0Unsafe RateLLaDA-Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LLaDA-InstructAttack=Zero-shot, Defense=Self-rem., Variant=Base2026.04 | 0 | 87.4 | |
| LLaDA-InstructAttack=DIJA, Defense=PPL, Variant=Base2026.04 | 0 | 90.8 | |
| LLaDA-InstructAttack=DIJA, Defense=PPL, Variant=+SAD2026.04 | 0 | 43.7 | |
| LLaDA-InstructAttack=DIJA, Defense=Self-rem., Variant=Base2026.04 | 0 | 90.8 | |
| LLaDA-InstructAttack=DIJA, Defense=Self-rem., Variant=+SAD2026.04 | 0 | 43.7 | |
| LLaDA-1.5Attack=DIJA, Defense=PPL, Variant=Base2026.04 | 0 | 90.6 | |
| LLaDA-1.5Attack=DIJA, Defense=PPL, Variant=+SAD2026.04 | 0 | 36.9 | |
| LLaDA-1.5Attack=DIJA, Defense=Self-rem., Variant=Base2026.04 | 0 | 90.6 | |
| Dream-InstructAttack=Zero-shot, Defense=None, Variant=+SAD2026.04 | 0 | 40.7 | |
| Dream-InstructAttack=Zero-shot, Defense=PPL, Variant=+SAD2026.04 | 0 | 40.7 | |
| Dream-InstructAttack=Zero-shot, Defense=DiffuGuard, Variant=Base2026.04 | 0 | 83.2 | |
| Dream-InstructAttack=Zero-shot, Defense=Self-rem.2026.04 | 0 | — | |
| Dream-Instruct + SADAttack=Zero-shot, Defense=Self-rem., Safety Scale (η)=2.02026.04 | 0 | — | |
| LLaDA-InstructAttack=DIJA, Defense=DiffuGuard, Variant=Base2026.04 | 0.2 | 90.8 | |
| LLaDA-InstructAttack=DIJA, Defense=DiffuGuard, Variant=+SAD2026.04 | 0.2 | 43.7 | |
| LLaDA-1.5Attack=Zero-shot, Defense=Self-rem., Variant=Base2026.04 | 0.2 | 83 | |
| LLaDA-1.5Attack=DIJA, Defense=None, Variant=Base2026.04 | 0.2 | 90.6 | |
| LLaDA-1.5Attack=DIJA, Defense=Self-rem., Variant=+SAD2026.04 | 0.2 | 36.9 | |
| Dream-InstructAttack=Zero-shot, Defense=None, Variant=Base2026.04 | 0.2 | 83.2 | |
| Dream-InstructAttack=Zero-shot, Defense=PPL, Variant=Base2026.04 | 0.2 | 83.2 | |
| Dream-InstructAttack=Zero-shot, Defense=Self-rem., Variant=Base2026.04 | 0.2 | 83.2 | |
| Dream-InstructAttack=Zero-shot, Defense=Self-rem., Variant=+SAD2026.04 | 0.2 | 40.7 | |
| Dream-InstructAttack=Zero-shot, Defense=PPL2026.04 | 0.2 | — | |
| Dream-Instruct + SADAttack=Zero-shot, Defense=PPL, Safety Scale (η)=2.02026.04 | 0.2 | — | |
| Dream-Instruct + SADAttack=Zero-shot, Defense=DiffuGuard, Safety Scale (η)=2.02026.04 | 0.2 | — | |
| LLaDA-InstructAttack=DIJA, Defense=None, Variant=Base2026.04 | 0.4 | 90.8 | |
| LLaDA-InstructAttack=DIJA, Defense=None, Variant=+SAD2026.04 | 0.4 | 43.7 | |
| Dream-InstructAttack=Zero-shot, Defense=DiffuGuard, Variant=+SAD2026.04 | 0.4 | 40.7 | |
| Dream-InstructAttack=Zero-shot, Defense=None2026.04 | 0.4 | 83.2 | |
| Dream-Instruct + SADAttack=Zero-shot, Defense=None, Safety Scale (η)=2.02026.04 | 0.4 | 40.7 | |
| Dream-InstructAttack=Zero-shot, Defense=DiffuGuard2026.04 | 0.4 | — | |
| LLaDA-1.5Attack=DIJA, Defense=DiffuGuard, Variant=Base2026.04 | 0.6 | 90.6 | |
| LLaDA-1.5Attack=Zero-shot, Defense=PPL, Variant=Base2026.04 | 0.8 | 83 | |
| LLaDA-1.5Attack=DIJA, Defense=None, Variant=+SAD2026.04 | 0.8 | 36.9 | |
| LLaDA-1.5Attack=DIJA, Defense=DiffuGuard, Variant=+SAD2026.04 | 0.8 | 36.9 | |
| LLaDA-InstructAttack=Zero-shot, Defense=Self-rem.2026.04 | 0.8 | — | |
| LLaDA-InstructAttack=Zero-shot, Defense=PPL, Variant=Base2026.04 | 1 | 87.4 | |
| LLaDA-InstructAttack=Zero-shot, Defense=DiffuGuard, Variant=Base2026.04 | 1 | 87.4 | |
| LLaDA-1.5Attack=Zero-shot, Defense=DiffuGuard, Variant=Base2026.04 | 1 | 83 | |
| LLaDA-1.5 + SADAttack=Zero-shot, Defense=Self-rem., Safety Scale (η)=4.02026.04 | 1 | — | |
| LLaDA-1.5Attack=Zero-shot, Defense=None, Variant=Base2026.04 | 1.2 | 83 | |
| LLaDA-Instruct + SADAttack=Zero-shot, Defense=Self-rem., Safety Scale (η)=2.02026.04 | 1.4 | — | |
| LLaDA-InstructAttack=Zero-shot, Defense=None, Variant=Base2026.04 | 1.6 | 87.4 | |
| LLaDA-1.5Attack=Zero-shot, Defense=Self-rem.2026.04 | 1.6 | — | |
| LLaDA-1.5 + SADAttack=Zero-shot, Defense=PPL, Safety Scale (η)=4.02026.04 | 3 | — | |
| Dream-Instruct + SADAttack=DIJA, Defense=DiffuGuard, Safety Scale (η)=2.02026.04 | 3.1 | — | |
| LLaDA-1.5 + SADAttack=Zero-shot, Defense=None, Safety Scale (η)=4.02026.04 | 3.2 | 6 | |
| LLaDA-Instruct + SADAttack=Zero-shot, Defense=DiffuGuard, Safety Scale (η)=2.02026.04 | 3.4 | — | |
| LLaDA-1.5 + SADAttack=Zero-shot, Defense=DiffuGuard, Safety Scale (η)=4.02026.04 | 3.4 | — | |
| LLaDA-1.5Attack=Zero-shot, Defense=Self-rem., Variant=+SAD2026.04 | 3.8 | 6 | |
| LLaDA-InstructAttack=Zero-shot, Defense=PPL2026.04 | 4 | — | |
| LLaDA-Instruct + SADAttack=Zero-shot, Defense=None, Safety Scale (η)=2.0, Sampling steps=64, SAD Time Window (C)=[0, 18], HarmBench negation set (N)=10002026.04 | 4.2 | 6.6 | |
| LLaDA-Instruct + SADAttack=Zero-shot, Defense=PPL, Safety Scale (η)=2.02026.04 | 4.2 | — | |
| LLaDA-InstructAttack=Zero-shot, Defense=DiffuGuard2026.04 | 4.2 | — | |
| LLaDA-InstructAttack=Zero-shot, Defense=None, Variant=+SAD2026.04 | 4.4 | 6.6 | |
| LLaDA-1.5Attack=Zero-shot, Defense=PPL, Variant=+SAD2026.04 | 4.4 | 6 | |
| SFTBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=Supervised Fine-Tuning2026.06 | 4.55 | — | |
| LLaDA-InstructAttack=Zero-shot, Defense=PPL, Variant=+SAD2026.04 | 4.6 | 6.6 | |
| LLaDA-InstructAttack=Zero-shot, Defense=DiffuGuard, Variant=+SAD2026.04 | 4.6 | 6.6 | |
| LLaDA-InstructAttack=Zero-shot, Defense=None, Sampling steps=64, SAD Time Window (C)=[0, 18], HarmBench negation set (N)=10002026.04 | 4.6 | 87.4 | |
| LLaDA-1.5Attack=Zero-shot, Defense=PPL2026.04 | 4.6 | — | |
| AlphaAlignBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=RL-based alignment2026.06 | 4.75 | — | |
| PolicyAlignBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=Policy-Based Safety Alignment2026.06 | 4.75 | — | |
| LLaDA-1.5Attack=Zero-shot, Defense=None, Variant=+SAD2026.04 | 4.8 | 6 | |
| NSPOBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=Null-space constrained policy optimization2026.06 | 4.89 | — | |
| GRPO+PolicyBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=Group Relative Policy Optimization with policy-conditioned reward2026.06 | 4.98 | — | |
| LLaDA-1.5Attack=Zero-shot, Defense=DiffuGuard, Variant=+SAD2026.04 | 5 | 6 | |
| LLaDA-1.5Attack=Zero-shot, Defense=DiffuGuard2026.04 | 5 | — | |
| LLaDA-InstructAttack=Zero-shot, Defense=Self-rem., Variant=+SAD2026.04 | 5.2 | 6.6 | |
| LLaDA-1.5Attack=Zero-shot, Defense=None2026.04 | 5.2 | 83 | |
| Dream-InstructAttack=DIJA, Defense=DiffuGuard2026.04 | 6.4 | — | |
| LLaDA-1.5 + SADAttack=PAD, Defense=Self-rem., Safety Scale (η)=4.02026.04 | 7.2 | — | |
| Dream-Instruct + SADAttack=PAD, Defense=Self-rem., Safety Scale (η)=2.02026.04 | 8 | — | |
| LLaDA-1.5Attack=PAD, Defense=Self-rem., Variant=+SAD2026.04 | 8.2 | 0.6 | |
| Dream-InstructAttack=PAD, Defense=Self-rem.2026.04 | 8.8 | — | |
| LLaDA-1.5Attack=PAD, Defense=Self-rem., Variant=Base2026.04 | 9 | 0.2 | |
| SFTBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=Supervised Fine-Tuning2026.06 | 9.5 | — | |
| Dream-InstructAttack=DIJA, Defense=Self-rem., Variant=+SAD2026.04 | 10 | 24.1 | |
| LLaDA-1.5Attack=PAD, Defense=Self-rem.2026.04 | 10.8 | — | |
| Dream-InstructAttack=PAD, Defense=Self-rem., Variant=+SAD2026.04 | 11 | 31.8 | |
| PolicyAlignBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=Policy-Based Safety Alignment2026.06 | 11.18 | — | |
| LLaDA-InstructAttack=PAD, Defense=Self-rem., Variant=Base2026.04 | 11.6 | 0.6 | |
| AlphaAlignBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=RL-based alignment2026.06 | 11.63 | — | |
| Dream-Instruct + SADAttack=DIJA, Defense=PPL, Safety Scale (η)=2.02026.04 | 11.7 | — | |
| LLaDA-Instruct + SADAttack=PAD, Defense=Self-rem., Safety Scale (η)=2.02026.04 | 11.8 | — | |
| LLaDA-Instruct + SADAttack=DIJA, Defense=PPL, Safety Scale (η)=2.02026.04 | 11.9 | — | |
| NSPOBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=Null-space constrained policy optimization2026.06 | 11.9 | — | |
| LLaDA-InstructAttack=DIJA, Defense=PPL2026.04 | 12.2 | — | |
| GRPO+PolicyBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=Group Relative Policy Optimization with policy-conditioned reward2026.06 | 12.26 | — | |
| PolicyAlignBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=Policy-Based Safety Alignment2026.06 | 12.4 | — | |
| ICLBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=In-Context Learning (System Prompt)2026.06 | 12.59 | — | |
| Dream-InstructAttack=DIJA, Defense=PPL2026.04 | 13.2 | — | |
| GRPO+PolicyBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=Group Relative Policy Optimization with policy-conditioned reward2026.06 | 13.22 | — | |
| AlphaAlignBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=RL-based alignment2026.06 | 13.57 | — | |
| SFTBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=Supervised Fine-Tuning2026.06 | 13.67 | — | |
| NSPOBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=Null-space constrained policy optimization2026.06 | 13.67 | — | |
| Dream-InstructAttack=PAD, Defense=Self-rem., Variant=Base2026.04 | 13.8 | 33.8 | |
| Dream-Instruct + SADAttack=PAD, Defense=PPL, Safety Scale (η)=2.02026.04 | 14 | — | |
| LLaDA-1.5Attack=DIJA, Defense=PPL2026.04 | 14.3 | — | |
| LLaDA-1.5 + SADAttack=DIJA, Defense=PPL, Safety Scale (η)=4.02026.04 | 14.3 | — |