Loading the SOTA2 catalog…
Contrastive Preference Learning: Learning from Human Feedback without RL · SOTA2 Research