Loading the SOTA2 catalog…
Inverse Preference Learning: Preference-based RL without a Reward Function · SOTA2 Research