Loading the SOTA2 catalog…
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment · SOTA2 Research