Loading the SOTA2 catalog…
Selective Preference Optimization via Token-Level Reward Function Estimation · SOTA2 Research