Top-k Selection on Top-k synthetic degree-8 Gaussian linear
1.68Final Cumulative RegretDFHPG
Evaluation Results
| Method | Links | |
|---|---|---|
| DFHPGFeedback regime=Pure bandit feedback, Horizon (T)=2,000, Replications=302026.05 | 1.68 | |
| DFHPG-0Feedback regime=Pure bandit feedback, Horizon (T)=2,000, Replications=302026.05 | 1.99 | |
| DFHPG-1Feedback regime=Pure bandit feedback, Horizon (T)=2,000, Replications=302026.05 | 4 | |
| DFHPG-0Feedback regime=Semi-bandit feedback, Horizon (T)=2,000, Replications=302026.05 | 6.65 | |
| TSCBFeedback regime=Pure bandit feedback, Horizon (T)=2,000, Replications=302026.05 | 6.93 | |
| GREEDYCBFeedback regime=Pure bandit feedback, Horizon (T)=2,000, Replications=302026.05 | 6.94 | |
| DFHPGFeedback regime=Semi-bandit feedback, Horizon (T)=2,000, Replications=302026.05 | 7.1 | |
| ϵ-GREEDYCBFeedback regime=Pure bandit feedback, Horizon (T)=2,000, Replications=302026.05 | 7.17 |