Harmfulness Evaluation on AdvBench
1.04Harmfulness ScoreBase
Evaluation Results
| Method | Links | |
|---|---|---|
| Base2025.10 | 1.04 | |
| Base2025.10 | 1.04 | |
| SAFEPATH-FTModel=R1-Llama-8B2025.08 | 1.06 | |
| ThinkingIModel=R1-Qwen-7B2025.08 | 1.08 | |
| ThinkingIModel=R1-Llama-8B2025.08 | 1.14 | |
| SAFEPATH-FTModel=R1-Qwen-7B2025.08 | 1.16 | |
| SafeKeyModel=R1-Llama-8B2025.08 | 1.34 | |
| ReasoningGuardModel=R1-Llama-8B2025.08 | 1.4 | |
| RealSafe-R1Model=R1-Qwen-7B2025.08 | 1.42 | |
| SafeKeyModel=R1-Qwen-7B2025.08 | 1.42 | |
| SAFEPATH-ZSModel=R1-Llama-8B2025.08 | 1.48 | |
| RealSafe-R1Model=R1-Llama-8B2025.08 | 1.52 | |
| ReasoningGuardModel=R1-Qwen-7B2025.08 | 1.78 | |
| SAFEPATH-ZSModel=R1-Qwen-7B2025.08 | 2.1 | |
| Self-ReminderModel=R1-Qwen-7B2025.08 | 2.32 | |
| Self-ReminderModel=R1-Llama-8B2025.08 | 2.32 | |
| ParaphraseModel=R1-Llama-8B2025.08 | 2.92 | |
| SmoothLLMModel=R1-Llama-8B2025.08 | 3.1 | |
| ParaphraseModel=R1-Qwen-7B2025.08 | 3.34 | |
| No DefenseModel=R1-Llama-8B2025.08 | 3.54 | |
| SmoothLLMModel=R1-Qwen-7B2025.08 | 3.58 | |
| SafeDecodingModel=R1-Llama-8B2025.08 | 3.64 | |
| No DefenseModel=R1-Qwen-7B2025.08 | 3.74 | |
| SafeDecodingModel=R1-Qwen-7B2025.08 | 3.74 | |
| SFT2025.10 | 4.87 | |
| SFT2025.10 | 4.87 | |
| RLVR2025.10 | 4.99 | |
| RLVR2025.10 | 4.99 |