ResearchTasksEnd-to-end policy optimizationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedSequential cooperative bandit testbed (test)CAPO10.2Mean Normalized Regret AUC20Apr 21, 2026