Loading the SOTA2 catalog…
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training · SOTA2 Research