Linear off-policy prediction on Two-state environment
1.697Max RMSEGTD2
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GTD2alpha=0.01, total runs=102026.05 | 1.697 | 0 | |
| ETDalpha=0.01, total runs=102026.05 | 1.72 | 0 | |
| RETDalpha=0.01, total runs=102026.05 | 1.754 | 0 | |
| TETDalpha=0.01, total runs=102026.05 | 1.809 | 0 | |
| TDRCalpha=0.01, total runs=102026.05 | 1.916 | 0 | |
| CETDalpha=0.01, total runs=102026.05 | 2.043 | 0 | |
| TDCalpha=0.01, total runs=102026.05 | 2.659 | 0 | |
| TDalpha=0.01, total runs=102026.05 | 7.1 | 10 |