Optimal Policy Estimation on Continuous Simulation Setting (epsilon = 0.5)
0.11Mean RegretSZonly
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SZonlySample size (n)=1000, Number of replications=50, Gaussian kernel bandwidth selection=median heuristic, Penalty tuning=cross-validation, Projection step estimation method=linear regression2022.09 | 0.11 | 0.0018 | |
| SuperSample size (n)=1000, Number of replications=50, Gaussian kernel bandwidth selection=median heuristic, Penalty tuning=cross-validation, Projection step estimation method=linear regression2022.09 | 0.11 | 0.0018 | |
| SonlySample size (n)=1000, Number of replications=50, Gaussian kernel bandwidth selection=median heuristic, Penalty tuning=cross-validation, Projection step estimation method=linear regression2022.09 | 0.4 | 0.0023 |