Harmlessness evaluation on HH-RLHF (test)
83.33Win RateAPL
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| APLModel=Mistral-7B2025.05 | 83.33 | — | 0.43 | 2.6 | |
| APL (w/o Adv)Model=Mistral-7B2025.05 | 80 | — | 1.3 | 1.9 | |
| APL (RM)Model=Mistral-7B2025.05 | 76.67 | — | 0.69 | 2.28 | |
| DPOModel=Mistral-7B2025.05 | 71.67 | — | 2.12 | 1.38 | |
| CAPOModel=Mistral-7B2025.05 | 70 | — | 1.9 | 1.46 | |
| APLModel=Llama-3-8B2025.05 | 56.67 | — | 1.34 | 2.09 | |
| APL (w/o Adv)Model=Llama-3-8B2025.05 | 54.17 | — | 1.69 | 1.97 | |
| DPOModel=Llama-3-8B2025.05 | 52.5 | — | 1.34 | 1.99 | |
| CAPOModel=Llama-3-8B2025.05 | 51 | — | 2.08 | 2.05 | |
| APL (RM)Model=Llama-3-8B2025.05 | 50.83 | — | 1.56 | 1.9 | |
| BaseModel=Mistral-7B2025.05 | 50 | — | 5.88 | 0.38 | |
| BaseModel=Llama-3-8B2025.05 | 50 | — | 1.95 | 2.02 | |
| BaseBackbone=MPT-7B-Instruct2024.02 | — | 40 | — | — | |
| DeALBackbone=MPT-7B-Instruct, Reward Model=R_harmless2024.02 | — | 57 | — | — | |
| DeALBackbone=MPT-7B-Instruct, Reward Model=R_helpful2024.02 | — | 37 | — | — | |
| DeALBackbone=MPT-7B-Instruct, Reward Model=R_hh2024.02 | — | 67 | — | — | |
| Harmless rerankBackbone=MPT-7B-Instruct2024.02 | — | 47 | — | — | |
| Helpful rerankBackbone=MPT-7B-Instruct2024.02 | — | 40 | — | — | |
| pa (for safety)Backbone=MPT-7B-Instruct2024.02 | — | 43 | — | — |