Loading the SOTA2 catalog…
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · SOTA2 Research