Loading the SOTA2 catalog…
RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs · SOTA2 Research