Safety Evaluation on ASB, InjecAgent, AHarm, and ASecBench
61.1ASB ScoreQwen3-8B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Qwen3-8BBase Model=Qwen3-8B, Training Strategy=Base2026.06 | 61.1 | 5.5 | 52.9 | 91.3 | 52.7 | |
| Qwen3-14BBase Model=Qwen3-14B, Training Strategy=Base2026.06 | 59.3 | 2.9 | 57.8 | 99 | 54.8 | |
| GLM-4.7-FlashBase Model=GLM-4.7-Flash, Training Strategy=Base2026.06 | 54.7 | 0.4 | 23.6 | 69.5 | 37.1 | |
| DPOBase Model=Qwen3-8B, Training Strategy=+ DPO2026.06 | 52.8 | 4 | 6.2 | 87.5 | 37.6 | |
| DPOBase Model=Qwen3-14B, Training Strategy=+ DPO, Training Status=Failure (†)2026.06 | 45.6 | 0.6 | 0 | 1 | 11.8 | |
| SFTBase Model=Qwen3-8B, Training Strategy=+ SFT2026.06 | 43.8 | 21.9 | 8.1 | 82.5 | 39.1 | |
| RUBASBase Model=GLM-4.7-Flash, Training Strategy=+ RUBAS2026.06 | 40.4 | 0 | 2.5 | 69.8 | 28.2 | |
| SFTBase Model=Qwen3-14B, Training Strategy=+ SFT2026.06 | 38.5 | 10.2 | 8.3 | 95.3 | 38.1 | |
| RuleBase Model=GLM-4.7-Flash, Training Strategy=+ Rule, Training Status=Failure (†)2026.06 | 32.3 | 0 | 1.1 | 64.8 | 24.6 | |
| SFTBase Model=GLM-4.7-Flash, Training Strategy=+ SFT2026.06 | 32.2 | 36 | 1.3 | 77.5 | 36.8 | |
| GuardModelBase Model=Qwen3-8B, Training Strategy=+ GuardModel2026.06 | 31.8 | 2.8 | 2.2 | 67 | 26 | |
| GuardModelBase Model=GLM-4.7-Flash, Training Strategy=+ GuardModel, Training Status=Failure (†)2026.06 | 29.9 | 0 | 2.5 | 50.5 | 20.7 | |
| RuleBase Model=Qwen3-8B, Training Strategy=+ Rule2026.06 | 29.4 | 2.3 | 0 | 65 | 24.2 | |
| DPOBase Model=GLM-4.7-Flash, Training Strategy=+ DPO, Training Status=Failure (†)2026.06 | 27.7 | 0 | 0 | 0 | 6.9 | |
| GuardModelBase Model=Qwen3-14B, Training Strategy=+ GuardModel2026.06 | 27.1 | 0.2 | 2.6 | 81 | 27.7 | |
| RuleBase Model=Qwen3-14B, Training Strategy=+ Rule2026.06 | 26.1 | 0.7 | 1 | 77.8 | 26.4 | |
| RUBASBase Model=Qwen3-14B, Training Strategy=+ RUBAS2026.06 | 21 | 0 | 0 | 78 | 24.8 | |
| RUBASBase Model=Qwen3-8B, Training Strategy=+ RUBAS2026.06 | 16.3 | 0.1 | 0 | 47.3 | 15.9 |