Loading the SOTA2 catalog…
Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight Guarantees · SOTA2 Research