Reinforcement Learning on Grid World Npick=5, Sparse (test)
0.7Maximum Average ReturnLOVE
Evaluation Results
| Method | Links | |
|---|---|---|
| LOVE2022.12 | 0.7 | |
| Option-Criticnumber of options=82022.12 | 0 |
| Method | Links | |
|---|---|---|
| LOVE2022.12 | 0.7 | |
| Option-Criticnumber of options=82022.12 | 0 |