Loading the SOTA2 catalog…
Intentionally-underestimated Value Function at Terminal State for Temporal-difference Learning with Mis-designed Reward · SOTA2 Research