Harmfulness Evaluation on AutoRAN
1.32Harmfulness ScoreSAFEPATH-FT
Evaluation Results
| Method | Links | |
|---|---|---|
| SAFEPATH-FTModel=R1-Llama-8B2025.08 | 1.32 | |
| RealSafe-R1Model=R1-Qwen-7B2025.08 | 1.44 | |
| RealSafe-R1Model=R1-Llama-8B2025.08 | 1.8 | |
| SAFEPATH-FTModel=R1-Qwen-7B2025.08 | 1.88 | |
| ReasoningGuardModel=R1-Qwen-7B2025.08 | 1.92 | |
| SAFEPATH-ZSModel=R1-Llama-8B2025.08 | 2.08 | |
| ThinkingIModel=R1-Qwen-7B2025.08 | 2.12 | |
| ReasoningGuardModel=R1-Llama-8B2025.08 | 2.2 | |
| ThinkingIModel=R1-Llama-8B2025.08 | 2.28 | |
| SAFEPATH-ZSModel=R1-Qwen-7B2025.08 | 2.42 | |
| SafeKeyModel=R1-Qwen-7B2025.08 | 2.56 | |
| Self-ReminderModel=R1-Qwen-7B2025.08 | 2.78 | |
| SafeKeyModel=R1-Llama-8B2025.08 | 2.9 | |
| No DefenseModel=R1-Qwen-7B2025.08 | 3.2 | |
| SafeDecodingModel=R1-Qwen-7B2025.08 | 3.22 | |
| SmoothLLMModel=R1-Llama-8B2025.08 | 3.22 | |
| ParaphraseModel=R1-Llama-8B2025.08 | 3.28 | |
| ParaphraseModel=R1-Qwen-7B2025.08 | 3.3 | |
| No DefenseModel=R1-Llama-8B2025.08 | 3.34 | |
| SafeDecodingModel=R1-Llama-8B2025.08 | 3.36 | |
| Self-ReminderModel=R1-Llama-8B2025.08 | 3.38 | |
| SmoothLLMModel=R1-Qwen-7B2025.08 | 3.44 |