Safety Evaluation on Trotter
30.8OverRefusal ScoreBase
Evaluation Results
| Method | Links | |
|---|---|---|
| BaseModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 30.8 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 28.8 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 27.3 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 26.3 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 23.7 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 23.7 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 23.7 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 21.7 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 21.7 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 18.7 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 14.4 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 11.1 |