Offline Reinforcement Learning on D4RL Gym walker2d (full-replay)
108.4Normalized ReturnBCQ
Evaluation Results
| Method | Links | |
|---|---|---|
| BCQImplementation=Reproduced2021.10 | 108.4 | |
| SPQRBase algorithm=EDAC2024.01 | 102.2 | |
| SPQRBase algorithm=SAC-Min2024.01 | 101.1 | |
| EDACImplementation=Ours2021.10 | 99.8 | |
| EDAC2024.01 | 99.8 | |
| CQL-Min2024.01 | 98.96 | |
| MORELImplementation=Reproduced2021.10 | 96.9 | |
| SAC-NImplementation=Ours2021.10 | 94.6 | |
| SAC-Min2024.01 | 94.6 | |
| CQLImplementation=Reproduced2021.10 | 94.2 | |
| BRACImplementation=Reproduced2021.10 | 71 | |
| BCImplementation=Reproduced2021.10 | 68.8 | |
| BC2024.01 | 68.8 | |
| UWACImplementation=Reproduced2021.10 | 60.7 | |
| SACImplementation=Reproduced2021.10 | 27.9 | |
| BEARImplementation=Reproduced2021.10 | 18.9 | |
| REMImplementation=Reproduced2021.10 | 1.3 |