Loading the SOTA2 catalog…
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL · SOTA2 Research