Loading the SOTA2 catalog…
Listening to the Echo: User-Reaction Aware Policy Optimization via Scalar-Verbal Hybrid Reinforcement Learning · SOTA2 Research