Offline Reinforcement Learning on D4RL Gym walker2d (expert)
121.6Normalized Avg ReturnCQL-Min
Evaluation Results
| Method | Links | |
|---|---|---|
| CQL-Min2024.01 | 121.6 | |
| SAC-Min2024.01 | 116.7 | |
| RORLImplementation Source=Ours, Ensemble Size=102022.06 | 115.4 | |
| SPQRBase algorithm=SAC-Min2024.01 | 115.2 | |
| EDACImplementation=Ours2021.10 | 115.1 | |
| EDACImplementation Source=Paper2022.06 | 115.1 | |
| EDAC2024.01 | 115.1 | |
| CQLImplementation=Reproduced2021.10 | 109.3 | |
| CQL2022.06 | 109.3 | |
| BCImplementation=Reproduced2021.10 | 108.7 | |
| BC2022.06 | 108.7 | |
| BC2024.01 | 108.7 | |
| PBRL2022.06 | 108.3 | |
| BEARImplementation=Reproduced2021.10 | 107.7 | |
| SAC-NImplementation=Ours2021.10 | 107.4 | |
| CQLImplementation=Original Paper2021.10 | 107 | |
| BCQImplementation=Reproduced2021.10 | 106.3 | |
| EDAC-10Implementation Source=Reproduced, Ensemble Size=102022.06 | 57.8 | |
| MORELImplementation=Reproduced2021.10 | 40.1 | |
| BRACImplementation=Reproduced2021.10 | 12.2 | |
| UWACImplementation=Reproduced2021.10 | 1.5 | |
| SAC-10Implementation Source=Reproduced, Ensemble Size=102022.06 | 1.2 | |
| SACImplementation=Reproduced2021.10 | 0.7 | |
| REMImplementation=Reproduced2021.10 | 0.7 |