ResearchTasksMulti-armed bandit policy evaluationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedSkew-t reward distribution Δ = 0.05 1,000 MC replicates syntheticUCB191.49Coverage (Arm1)6Feb 26, 2026Student-t reward distribution Δ = 0.05 synthetic (1,000 MC replicates)CP-Bandit81.29Coverage (Arm1)6Feb 26, 2026Gaussian reward distribution (Δ = 0.05) 1,000 MC replicates syntheticUCB184.01Coverage (Arm1)6Feb 26, 2026
Skew-t reward distribution Δ = 0.05 1,000 MC replicates syntheticUCB191.49Coverage (Arm1)6Feb 26, 2026
Student-t reward distribution Δ = 0.05 synthetic (1,000 MC replicates)CP-Bandit81.29Coverage (Arm1)6Feb 26, 2026
Gaussian reward distribution (Δ = 0.05) 1,000 MC replicates syntheticUCB184.01Coverage (Arm1)6Feb 26, 2026