Loading the SOTA2 catalog…
Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models · SOTA2 Research