Safety Alignment on SafeRLHF
83Win RateChain-of-Thought
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Chain-of-ThoughtMain Model=Vicuna, Inference Strategy=Chain-of-Thought2025.06 | 83 | 8 | 9 | |
| Best-of-NMain Model=Vicuna, Inference Strategy=Best-of-N2025.06 | 77 | 10 | 13 | |
| Multi-Agent DebateMain Model=Vicuna, Inference Strategy=Multi-Agent Debate2025.06 | 76 | 10 | 14 | |
| RLAIFtraining_context=after re-alignment training2025.06 | 72 | 22 | 6 | |
| Chain-of-ThoughtMain Model=GPT-4o-mini, Projector=GPT-4o-mini2025.06 | 71 | 6 | 23 | |
| Best-of-NMain Model=GPT-4o-mini, Projector=GPT-4o-mini2025.06 | 62 | 20 | 18 | |
| Multi-Agent DebateMain Model=GPT-4o-mini, Projector=GPT-4o-mini2025.06 | 61 | 21 | 18 | |
| Self-Refine (Debate)training_context=after re-alignment training2025.06 | 57 | 32 | 11 |