Loading the SOTA2 catalog…
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game · SOTA2 Research