Loading the SOTA2 catalog…
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization · SOTA2 Research