Loading the SOTA2 catalog…
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs · SOTA2 Research