Loading the SOTA2 catalog…
S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning · SOTA2 Research