Loading the SOTA2 catalog…
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback · SOTA2 Research