Loading the SOTA2 catalog…
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates · SOTA2 Research