Safety Evaluation on XSTest (OverRefusal Score)
99.6OverRefusal ScoreSafeChain
Evaluation Results
| Method | Links | |
|---|---|---|
| SafeChainModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 99.6 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 99.2 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 99.2 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 98.8 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 98.4 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 96.4 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 96 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 94.4 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 94.3 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 93.6 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 93.6 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 89.6 |