Safety Evaluation on HEX-PHI (test)
2Harmfulness Score (Llama-Guard-3B)Original
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OriginalModel=Llama-2-7B-Chat, Utility Context=GSM8K2025.03 | 2 | — | — | 1.5 | |
| OriginalModel=Llama-2-7B-Chat, Utility Context=PubMedQA2025.03 | 2 | — | — | 1.5 | |
| RESTAModel=Llama-2-7B-Chat, Utility Context=PubMedQA2025.03 | 4.2 | — | — | 3.8 | |
| RESTAModel=Llama-2-7B-Chat, Utility Context=GSM8K2025.03 | 4.3 | — | — | 4.1 | |
| SafeMERGEModel=Llama-2-7B-Chat, Utility Context=PubMedQA2025.03 | 4.3 | — | — | 4.1 | |
| RESTA-InstructModel=Llama-2-7B-Chat, Utility Context=PubMedQA2025.03 | 5.3 | — | — | 5 | |
| SafeMERGEModel=Llama-2-7B-Chat, Utility Context=GSM8K2025.03 | 5.7 | — | — | 5.2 | |
| SafeLoRAModel=Llama-2-7B-Chat, Utility Context=PubMedQA2025.03 | 5.9 | — | — | 5.6 | |
| SafeInstructModel=Qwen-2-7B-Instruct, Utility Context=PubMedQA2025.03 | 5.9 | — | — | 5.3 | |
| SafeMERGEModel=Qwen-2-7B-Instruct, Utility Context=PubMedQA2025.03 | 5.9 | — | — | 5.2 | |
| SafeInstructModel=Llama-2-7B-Chat, Utility Context=GSM8K2025.03 | 6.2 | — | — | 5.9 | |
| Fine-tunedModel=Llama-2-7B-Chat, Utility Context=PubMedQA2025.03 | 6.2 | — | — | 5.9 | |
| SafeInstructModel=Llama-2-7B-Chat, Utility Context=PubMedQA2025.03 | 6.3 | — | — | 5.9 | |
| SafeMERGEModel=Llama-3.1-8B-Instruct, Utility Context=GSM8K2025.03 | 6.3 | — | — | 6.1 | |
| RESTA-InstructModel=Llama-2-7B-Chat, Utility Context=GSM8K2025.03 | 6.8 | — | — | 6.3 | |
| SafeMERGEModel=Llama-3.1-8B-Instruct, Utility Context=PubMedQA2025.03 | 6.8 | — | — | 6.4 | |
| SafeLoRAModel=Llama-2-7B-Chat, Utility Context=GSM8K2025.03 | 6.9 | — | — | 6.4 | |
| RESTAModel=Llama-3.1-8B-Instruct, Utility Context=GSM8K2025.03 | 6.9 | — | — | 6.7 | |
| SafeLoRAModel=Llama-3.1-8B-Instruct, Utility Context=GSM8K2025.03 | 7.1 | — | — | 6.8 | |
| RESTAModel=Llama-3.1-8B-Instruct, Utility Context=PubMedQA2025.03 | 7.1 | — | — | 6.7 | |
| SafeInstructModel=Llama-3.1-8B-Instruct, Utility Context=GSM8K2025.03 | 7.2 | — | — | 6.8 | |
| RESTA-InstructModel=Llama-3.1-8B-Instruct, Utility Context=GSM8K2025.03 | 7.2 | — | — | 6.8 | |
| SafeMERGEModel=Qwen-2-7B-Instruct, Utility Context=GSM8K2025.03 | 7.5 | — | — | 6.9 | |
| OriginalModel=Llama-3.1-8B-Instruct, Utility Context=GSM8K2025.03 | 7.9 | — | — | 7.7 | |
| OriginalModel=Llama-3.1-8B-Instruct, Utility Context=PubMedQA2025.03 | 7.9 | — | — | 7.7 | |
| RESTA-InstructModel=Llama-3.1-8B-Instruct, Utility Context=PubMedQA2025.03 | 8.7 | — | — | 8.5 | |
| SafeMERGEModel=Qwen-2.5-7B-Instruct, Utility Context=PubMedQA2025.03 | 8.7 | — | — | 8.4 | |
| SafeInstructModel=Qwen-2-7B-Instruct, Utility Context=GSM8K2025.03 | 9.5 | — | — | 9.1 | |
| SafeInstructModel=Qwen-2.5-7B-Instruct, Utility Context=PubMedQA2025.03 | 9.5 | — | — | 8.9 | |
| SafeLoRAModel=Llama-3.1-8B-Instruct, Utility Context=PubMedQA2025.03 | 9.6 | — | — | 9.2 | |
| SafeInstructModel=Llama-3.1-8B-Instruct, Utility Context=PubMedQA2025.03 | 9.7 | — | — | 9.3 | |
| SafeMERGEModel=Qwen-2.5-7B-Instruct, Utility Context=GSM8K2025.03 | 10.3 | — | — | 9.9 | |
| SafeInstructModel=Qwen-2.5-7B-Instruct, Utility Context=GSM8K2025.03 | 11.3 | — | — | 10.8 | |
| OriginalModel=Qwen-2-7B-Instruct, Utility Context=GSM8K2025.03 | 11.5 | — | — | 11.1 | |
| OriginalModel=Qwen-2-7B-Instruct, Utility Context=PubMedQA2025.03 | 11.5 | — | — | 11.1 | |
| Fine-tunedModel=Llama-3.1-8B-Instruct, Utility Context=PubMedQA2025.03 | 12.2 | — | — | 11.8 | |
| Fine-tunedModel=Qwen-2-7B-Instruct, Utility Context=PubMedQA2025.03 | 13.2 | — | — | 12.8 | |
| OriginalModel=Qwen-2.5-7B-Instruct, Utility Context=GSM8K2025.03 | 13.2 | — | — | 12.8 | |
| OriginalModel=Qwen-2.5-7B-Instruct, Utility Context=PubMedQA2025.03 | 13.2 | — | — | 12.8 | |
| SafeLoRAModel=Qwen-2.5-7B-Instruct, Utility Context=PubMedQA2025.03 | 13.3 | — | — | 12.7 | |
| RESTA-InstructModel=Qwen-2-7B-Instruct, Utility Context=GSM8K2025.03 | 14.1 | — | — | 13.7 | |
| SafeLoRAModel=Qwen-2.5-7B-Instruct, Utility Context=GSM8K2025.03 | 14.4 | — | — | 13.7 | |
| RESTA-InstructModel=Qwen-2-7B-Instruct, Utility Context=PubMedQA2025.03 | 14.5 | — | — | 13.9 | |
| SafeLoRAModel=Qwen-2-7B-Instruct, Utility Context=PubMedQA2025.03 | 14.5 | — | — | 13.9 | |
| Fine-tunedModel=Llama-3.1-8B-Instruct, Utility Context=GSM8K2025.03 | 14.7 | — | — | 14.3 | |
| SafeLoRAModel=Qwen-2-7B-Instruct, Utility Context=GSM8K2025.03 | 14.8 | — | — | 14.2 | |
| RESTAModel=Qwen-2-7B-Instruct, Utility Context=PubMedQA2025.03 | 14.8 | — | — | 14.4 | |
| RESTA-InstructModel=Qwen-2.5-7B-Instruct, Utility Context=PubMedQA2025.03 | 15.2 | — | — | 14.6 | |
| RESTA-InstructModel=Qwen-2.5-7B-Instruct, Utility Context=GSM8K2025.03 | 15.4 | — | — | 14.8 | |
| RESTAModel=Qwen-2.5-7B-Instruct, Utility Context=PubMedQA2025.03 | 15.6 | — | — | 14.5 | |
| RESTAModel=Qwen-2-7B-Instruct, Utility Context=GSM8K2025.03 | 15.8 | — | — | 15.3 | |
| RESTAModel=Qwen-2.5-7B-Instruct, Utility Context=GSM8K2025.03 | 16.2 | — | — | 15.7 | |
| Fine-tunedModel=Llama-2-7B-Chat, Utility Context=GSM8K2025.03 | 16.4 | — | — | 16 | |
| Fine-tunedModel=Qwen-2-7B-Instruct, Utility Context=GSM8K2025.03 | 16.8 | — | — | 16.1 | |
| Fine-tunedModel=Qwen-2.5-7B-Instruct, Utility Context=GSM8K2025.03 | 17.3 | — | — | 16.7 | |
| Fine-tunedModel=Qwen-2.5-7B-Instruct, Utility Context=PubMedQA2025.03 | 17.8 | — | — | 17.2 | |
| AdvBench AlignmentAlignment Dataset=AdvBench2025.02 | — | 1.42 | 13.03 | — | |
| AdvBench-R2J AlignmentAlignment Dataset=AdvBench-R2J2025.02 | — | 1.2 | 7.58 | — | |
| AWQBackbone=Gemma-7B-Instruct, Quantization=AWQ2026.01 | — | — | 8.667 | — | |
| AWQBackbone=Llama-3.1-8B-Instruct, Quantization=AWQ2026.01 | — | — | 9 | — | |
| AWQBackbone=Qwen-2.5-7B-Instruct, Quantization=AWQ2026.01 | — | — | 14 | — | |
| AWQ-trustBackbone=Gemma-7B-Instruct, Quantization=AWQ-trust2026.01 | — | — | 9.333 | — | |
| AWQ-trustBackbone=Llama-3.1-8B-Instruct, Quantization=AWQ-trust2026.01 | — | — | 7.333 | — | |
| AWQ-trustBackbone=Qwen-2.5-7B-Instruct, Quantization=AWQ-trust2026.01 | — | — | 10 | — | |
| Full PrecisionBackbone=Gemma-7B-Instruct, Quantization=Full Precision2026.01 | — | — | 9.333 | — | |
| Full PrecisionBackbone=Llama-3.1-8B-Instruct, Quantization=Full Precision2026.01 | — | — | 9.667 | — | |
| Full PrecisionBackbone=Qwen-2.5-7B-Instruct, Quantization=Full Precision2026.01 | — | — | 12.667 | — | |
| w/o SFTAlignment Dataset=None2025.02 | — | 3.3 | 58.48 | — |