Loading the SOTA2 catalog…
A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation · SOTA2 Research