Loading the SOTA2 catalog…
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment · SOTA2 Research