Safety Evaluation on WildChat (test)
69.85WildChat ScoreSFT-DPO + LoRA
Evaluation Results
| Method | Links | |
|---|---|---|
| SFT-DPO + LoRABase Model=Qwen2.5-7B-Instruct, Alignment=SFT-DPO, Strategy=LoRA2026.02 | 69.85 | |
| SFTBase Model=Qwen2.5-7B-Instruct, Alignment=SFT2026.02 | 64.2 | |
| SFT-DPOBase Model=Llama3.1-8B-Instruct, Alignment=SFT-DPO2026.02 | 59.4 | |
| SFTBase Model=Llama3.1-8B-Instruct, Alignment=SFT2026.02 | 50 | |
| DPO + OGPSABase Model=Qwen2.5-7B-Instruct, Alignment=DPO, Strategy=OGPSA2026.02 | 49.4 | |
| SFT + OGPSABase Model=Llama3.1-8B-Instruct, Alignment=SFT, Strategy=OGPSA2026.02 | 47 | |
| DPOBase Model=Llama3.1-8B-Instruct, Alignment=DPO2026.02 | 42.8 | |
| SFT + LoRABase Model=Llama3.1-8B-Instruct, Alignment=SFT, Strategy=LoRA2026.02 | 42.6 | |
| DPO + OGPSABase Model=Llama3.1-8B-Instruct, Alignment=DPO, Strategy=OGPSA2026.02 | 38.4 | |
| SFT + MergeBase Model=Llama3.1-8B-Instruct, Alignment=SFT, Strategy=Merge2026.02 | 31.2 | |
| Instruct BaselineBase Model=Qwen2.5-7B-Instruct2026.02 | 16 | |
| Instruct BaselineBase Model=Llama3.1-8B-Instruct2026.02 | 15.8 | |
| SFT + General DataBase Model=Llama3.1-8B-Instruct, Alignment=SFT, Strategy=General Data2026.02 | 14.4 |