Offline Reinforcement Learning on D4RL walker2d-medium v0
1,269ReturnGELATO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GELATOdistance=metric2021.02 | 1,269 | — | |
| MOPOvariant=bootstrap2021.02 | 518 | — | |
| MBPO2021.02 | 370 | — | |
| GELATOdistance=l22021.02 | 312 | — | |
| Imitation2021.02 | 193 | — | |
| SAC2021.02 | 27 | — | |
| BCAlgorithmic Template=One-step2021.06 | — | 70.2 | |
| Easy BCQAlgorithmic Template=One-step2021.06 | — | 86.9 | |
| Exp. WeightAlgorithmic Template=One-step2021.06 | — | 80.3 | |
| Fu et al.Algorithmic Template=Iterative2021.06 | — | 81.1 | |
| Rev. KL Reg.Algorithmic Template=One-step2021.06 | — | 85.6 |