Reward Hacking Mitigation on Length Bias OA Length 1.0 (Evaluation)
15DominanceIR3 Method C (Constrained)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| IR3 Method C (Constrained)type=Constrained2026.02 | 15 | 68.5 | |
| IR3 Method B (Adversarial)type=Adversarial2026.02 | 16 | 67.2 | |
| IR3 Method A (Clean RL)type=Clean RL2026.02 | 21 | 63.5 | |
| IR3 Method D (Distillation)type=Distillation2026.02 | 24 | 71.8 | |
| Length Penaltyalpha=5e-42026.02 | 32 | 56.2 | |
| KL Regularizationbeta=2e-22026.02 | 35 | 54.8 | |
| Reward Clippingc=42026.02 | 36 | 53.8 | |
| PPO Clippingepsilon=0.12026.02 | 38 | 52.5 | |
| PPO on R_proxy2026.02 | 42 | 50 |