Offline Reinforcement Learning on D4RL Adroit door v0 (cloned)
40Normalized Average ReturnFu et al.
Evaluation Results
| Method | Links | |
|---|---|---|
| Fu et al.Algorithmic Template=Iterative2021.06 | 40 | |
| Easy BCQAlgorithmic Template=One-step2021.06 | 40 | |
| Exp. WeightAlgorithmic Template=One-step2021.06 | 10 | |
| EDACImplementation Source=Ours2021.10 | 9.6 | |
| CQLImplementation Source=Paper2021.10 | 3.5 | |
| CQLImplementation Source=Reproduced2021.10 | 2.4 | |
| BC2021.10 | 0 | |
| BCAlgorithmic Template=One-step2021.06 | 0 | |
| Rev. KL Reg.Algorithmic Template=One-step2021.06 | 0 | |
| SAC2021.10 | -0.3 | |
| REM2021.10 | -0.3 | |
| SAC-NImplementation Source=Ours2021.10 | -0.3 |