Loading the SOTA2 catalog…
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training · SOTA2 Research