Reward Hacking Mitigation on Excessive HH Harmless 1.0 (Evaluation)
8.2Reference Error RateIR3 Method B (Adversarial)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| IR3 Method B (Adversarial)type=Adversarial2026.02 | 8.2 | 91.2 | |
| IR3 Method C (Constrained)type=Constrained2026.02 | 8.8 | 90.9 | |
| IR3 Method A (Clean RL)type=Clean RL2026.02 | 10.8 | 90.8 | |
| IR3 Method D (Distillation)type=Distillation2026.02 | 11.5 | 90.2 | |
| InfoRM2026.02 | 14.5 | 90.5 | |
| Length Penaltyalpha=5e-42026.02 | 16.8 | 89 | |
| KL Regularizationbeta=2e-22026.02 | 18.2 | 88.2 | |
| Reward Clippingc=42026.02 | 19.5 | 89.2 | |
| PPO Clippingepsilon=0.12026.02 | 20.1 | 89.5 | |
| PPO on R_proxy2026.02 | 23.4 | 89.8 |