Loading the SOTA2 catalog…
Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling · SOTA2 Research