Loading the SOTA2 catalog…
Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards · SOTA2 Research