ResearchDatasetsSequential cooperative bandit testbedFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsEnd-to-end policy optimizationSequential cooperative bandit testbed (test)10.2Mean Normalized Regret AUC20
End-to-end policy optimizationSequential cooperative bandit testbed (test)10.2Mean Normalized Regret AUC20