Loading the SOTA2 catalog…
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States · SOTA2 Research