Harmfulness Evaluation on GCG
1.16Harmfulness ScoreSAFEPATH-FT
Evaluation Results
| Method | Links | |
|---|---|---|
| SAFEPATH-FTModel=R1-Llama-8B2025.08 | 1.16 | |
| SAFEPATH-ZSModel=R1-Llama-8B2025.08 | 1.34 | |
| ReasoningGuardModel=R1-Llama-8B2025.08 | 1.34 | |
| SafeKeyModel=R1-Llama-8B2025.08 | 1.4 | |
| SAFEPATH-ZSModel=R1-Qwen-7B2025.08 | 1.42 | |
| RealSafe-R1Model=R1-Qwen-7B2025.08 | 1.44 | |
| SafeKeyModel=R1-Qwen-7B2025.08 | 1.44 | |
| ReasoningGuardModel=R1-Qwen-7B2025.08 | 1.46 | |
| SAFEPATH-FTModel=R1-Qwen-7B2025.08 | 1.54 | |
| RealSafe-R1Model=R1-Llama-8B2025.08 | 1.6 | |
| ThinkingIModel=R1-Qwen-7B2025.08 | 1.86 | |
| ThinkingIModel=R1-Llama-8B2025.08 | 1.96 | |
| Self-ReminderModel=R1-Llama-8B2025.08 | 1.98 | |
| Self-ReminderModel=R1-Qwen-7B2025.08 | 2.36 | |
| ParaphraseModel=R1-Llama-8B2025.08 | 2.56 | |
| ParaphraseModel=R1-Qwen-7B2025.08 | 2.86 | |
| SmoothLLMModel=R1-Llama-8B2025.08 | 3.08 | |
| SmoothLLMModel=R1-Qwen-7B2025.08 | 3.16 | |
| No DefenseModel=R1-Llama-8B2025.08 | 3.32 | |
| SafeDecodingModel=R1-Llama-8B2025.08 | 3.48 | |
| SafeDecodingModel=R1-Qwen-7B2025.08 | 3.82 | |
| No DefenseModel=R1-Qwen-7B2025.08 | 4.16 |