Loading the SOTA2 catalog…
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts · SOTA2 Research