Harmfulness Evaluation on Mousetrap
3.78Harmfulness ScoreSelf-Reminder
Evaluation Results
| Method | Links | |
|---|---|---|
| Self-ReminderModel=R1-Llama-8B2025.08 | 3.78 | |
| ParaphraseModel=R1-Llama-8B2025.08 | 3.52 | |
| No DefenseModel=R1-Llama-8B2025.08 | 3.2 | |
| SafeKeyModel=R1-Llama-8B2025.08 | 3.14 | |
| SafeDecodingModel=R1-Llama-8B2025.08 | 3.12 | |
| SAFEPATH-ZSModel=R1-Llama-8B2025.08 | 2.9 | |
| ParaphraseModel=R1-Qwen-7B2025.08 | 2.82 | |
| SafeKeyModel=R1-Qwen-7B2025.08 | 2.72 | |
| SmoothLLMModel=R1-Llama-8B2025.08 | 2.72 | |
| No DefenseModel=R1-Qwen-7B2025.08 | 2.7 | |
| SafeDecodingModel=R1-Qwen-7B2025.08 | 2.66 | |
| SmoothLLMModel=R1-Qwen-7B2025.08 | 2.48 | |
| Self-ReminderModel=R1-Qwen-7B2025.08 | 2.32 | |
| RealSafe-R1Model=R1-Qwen-7B2025.08 | 1.92 | |
| SAFEPATH-ZSModel=R1-Qwen-7B2025.08 | 1.8 | |
| ThinkingIModel=R1-Qwen-7B2025.08 | 1.78 | |
| ThinkingIModel=R1-Llama-8B2025.08 | 1.66 | |
| RealSafe-R1Model=R1-Llama-8B2025.08 | 1.46 | |
| ReasoningGuardModel=R1-Llama-8B2025.08 | 1.34 | |
| SAFEPATH-FTModel=R1-Qwen-7B2025.08 | 1.3 | |
| ReasoningGuardModel=R1-Qwen-7B2025.08 | 1.14 | |
| SAFEPATH-FTModel=R1-Llama-8B2025.08 | 1.02 |