Loading the SOTA2 catalog…
CARE-RFT: Confidence-Anchored Reinforcement Finetuning for Reliable Reasoning in Large Language Models · SOTA2 Research