Preference evaluation on Anthropic-SafeRLHF benchmark
33.7Win Rateπbias (rubric-based preference attack)
Evaluation Results
| Method | Links | |
|---|---|---|
| πbias (rubric-based preference attack)Comparison=πbias vs. πseed, Training setting=Benchmark-only (B), Evaluator model=LLaMA-3-8B, Decoding strategy=Best-of-42026.02 | 33.7 | |
| πbias (rubric-based preference attack)Comparison=πbias vs. πseed, Training setting=Benchmark+Target (BT), Evaluator model=LLaMA-3-8B, Decoding strategy=Best-of-42026.02 | 23.9 |