Loading the SOTA2 catalog…
Secrets of RLHF in Large Language Models Part II: Reward Modeling · SOTA2 Research