Loading the SOTA2 catalog…
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning · SOTA2 Research