Loading the SOTA2 catalog…
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning · SOTA2 Research