Safety Evaluation on Fortress (JailBreak Score)
2.8JailBreak ScoreSInternal
Evaluation Results
| Method | Links | |
|---|---|---|
| SInternalModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 2.8 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 4.8 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 5.2 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 7.8 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 8 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 11 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 12.2 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 14.4 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 14.8 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 17 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 21.4 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 25.2 |