Offline Reinforcement Learning on D4RL Gym hopper (full-replay)
109.1Normalized ReturnSPQR
Evaluation Results
| Method | Links | |
|---|---|---|
| SPQRBase algorithm=SAC-Min2024.01 | 109.1 | |
| EDAC2024.01 | 105.4 | |
| CQL-Min2024.01 | 104.85 | |
| EDACImplementation=Ours2021.10 | 102.9 | |
| SAC-Min2024.01 | 102.9 | |
| SAC-NImplementation=Ours2021.10 | 101.9 | |
| CQLImplementation=Reproduced2021.10 | 94.4 | |
| MORELImplementation=Reproduced2021.10 | 81.8 | |
| BRACImplementation=Reproduced2021.10 | 62.7 | |
| BCQImplementation=Reproduced2021.10 | 60.9 | |
| BEARImplementation=Reproduced2021.10 | 57.7 | |
| UWACImplementation=Reproduced2021.10 | 31.1 | |
| REMImplementation=Reproduced2021.10 | 27.5 | |
| BCImplementation=Reproduced2021.10 | 19.9 | |
| BC2024.01 | 19.9 | |
| SACImplementation=Reproduced2021.10 | 7.4 |