Loading the SOTA2 catalog…
When Policies Cannot Be Retrained: A Unified Closed-Form View of Post-Training Steering in Offline Reinforcement Learning · SOTA2 Research