Loading the SOTA2 catalog…
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning · SOTA2 Research