Loading the SOTA2 catalog…
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning · SOTA2 Research