Reward Hacking Mitigation on Synthetic Goodhart 1.0 (Evaluation)
4.38R_gIR3 Method B (Adversarial)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| IR3 Method B (Adversarial)type=Adversarial2026.02 | 4.38 | 0.41 | |
| IR3 Method C (Constrained)type=Constrained2026.02 | 4.35 | 0.45 | |
| IR3 Method A (Clean RL)type=Clean RL2026.02 | 4.21 | 0.62 | |
| IR3 Method D (Distillation)type=Distillation2026.02 | 4.12 | 0.71 | |
| InfoRM2026.02 | 3.95 | 1.08 | |
| Length Penaltyalpha=5e-42026.02 | 3.85 | 1.32 | |
| KL Regularizationbeta=2e-22026.02 | 3.78 | 1.42 | |
| Reward Clippingc=42026.02 | 3.72 | 1.48 | |
| PPO interface clippingepsilon=0.12026.02 | 3.68 | 1.55 | |
| PPO on R_proxy2026.02 | 3.52 | 1.86 |