Safety Evaluation on XSTest
98.4Safety ScoreBase
Evaluation Results
| Method | Links | |
|---|---|---|
| Base2025.09 | 98.4 | |
| Qwen2.5 Instruct (72B)Evaluation Source=HELM2025.01 | 97.9 | |
| Qwen3.5-9BParameters=9B, Variant=Thinking2026.05 | 97.6 | |
| GPT-4o (2024-05-13)Evaluation Source=HELM2025.01 | 97.3 | |
| DeepSeek-V3Evaluation Source=HELM, Risk Control System=Enabled2025.01 | 97.1 | |
| o1 (2024-12-17)Evaluation Source=HELM2025.01 | 97 | |
| Qwen3.5-4BParameters=4B, Variant=Thinking2026.05 | 96.8 | |
| Ministral-3-14BParameters=14B, Variant=Thinking2026.05 | 96.8 | |
| Claude-3.7-SonnetEvaluation Source=HELM2025.01 | 96.4 | |
| DeepSeek-R1Evaluation Source=HELM, Chain of Thought=visible, Risk Control System=Enabled2025.01 | 95.3 | |
| DeepSeek-R1Evaluation Source=HELM, Chain of Thought=hidden, Risk Control System=Enabled2025.01 | 94.4 | |
| IQuest-Coder-V1-40B-ThinkingParameters=40B, Type=Thinking2026.03 | 94.3 | |
| OLMo-3-7BParameters=7B, Variant=Thinking2026.05 | 93.2 | |
| Mellum 2 (SFT)Post-training Stage=SFT, Parameters=2.5B/12B, Variant=Thinking2026.05 | 90.8 | |
| Qwen2.5-Coder-32B-InstructParameters=32B, Type=Instruct2026.03 | 90.6 | |
| Qwen3-Coder-480B-A35B-InstructParameters=480B-A35B, Type=Instruct2026.03 | 90.1 | |
| Mellum 2 (RL)Post-training Stage=RL, Parameters=2.5B/12B, Variant=Thinking2026.05 | 89.6 | |
| IQuest-Coder-V1-40B-InstructParameters=40B, Type=Instruct2026.03 | 89.3 | |
| Self-Improving PretrainingPre-training Data=RedPajama, Pre-training Strategy=Self-Improving2026.01 | 88.4 | |
| TARS2025.09 | 88.3 | |
| Llama Pretrain BaselinePre-training Data=RedPajama, Pre-training Strategy=Standard next token prediction2026.01 | 87.6 | |
| GRPO2025.09 | 86.8 | |
| Llama BasePre-training Strategy=Standard next token prediction2026.01 | 85.2 | |
| BackTrack2025.09 | 80 | |
| IPO2025.09 | 80 | |
| STAR2025.09 | 76.9 | |
| Self-Improving PretrainingPre-training Data=RedPajama, Pre-training Strategy=Self-Improving2026.01 | 49 | |
| Llama BasePre-training Strategy=Standard next token prediction2026.01 | 39.5 | |
| Llama Pretrain BaselinePre-training Data=RedPajama, Pre-training Strategy=Standard next token prediction2026.01 | 35 | |
| PCA-HMMModel=Mistral-7B, Diagnostic Method=PCA-HMM, Input Scope=User-window2026.05 | 15 | |
| PCA-HMMModel=Llama-3.1-8B, Diagnostic Method=PCA-HMM, Input Scope=User-window2026.05 | 9.5 | |
| PCA-HMMModel=OLMo3-7B, Diagnostic Method=PCA-HMM, Input Scope=User-window2026.05 | 6 |