Loading the SOTA2 catalog…
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification · SOTA2 Research