Loading the SOTA2 catalog…
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models · SOTA2 Research