Loading the SOTA2 catalog…
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts · SOTA2 Research