Loading the SOTA2 catalog…
R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging · SOTA2 Research