Loading the SOTA2 catalog…
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization · SOTA2 Research