Loading the SOTA2 catalog…
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning · SOTA2 Research