Loading the SOTA2 catalog…
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models · SOTA2 Research