Loading the SOTA2 catalog…
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation · SOTA2 Research