Loading the SOTA2 catalog…
MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement Learning · SOTA2 Research