Loading the SOTA2 catalog…
SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment · SOTA2 Research