Optimal Policy Estimation on Continuous Simulation Setting epsilon = 0.9
0.06Mean RegretSuper
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SuperSample size (n)=1000, Number of replications=50, Gaussian kernel bandwidth selection=median heuristic, Penalty tuning=cross-validation, Projection step estimation method=linear regression2022.09 | 0.06 | 0.0063 | |
| SZonlySample size (n)=1000, Number of replications=50, Gaussian kernel bandwidth selection=median heuristic, Penalty tuning=cross-validation, Projection step estimation method=linear regression2022.09 | 0.12 | 0.0529 | |
| SonlySample size (n)=1000, Number of replications=50, Gaussian kernel bandwidth selection=median heuristic, Penalty tuning=cross-validation, Projection step estimation method=linear regression2022.09 | 0.4 | 0.002 |