Loading the SOTA2 catalog…
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation · SOTA2 Research