Offline Reinforcement Learning on D4RL Gym halfcheetah (full-replay)
88.5Normalized ReturnSPQR
Evaluation Results
| Method | Links | |
|---|---|---|
| SPQRBase algorithm=SAC-Min2024.01 | 88.5 | |
| SACImplementation=Reproduced2021.10 | 86.8 | |
| SPQRBase algorithm=EDAC2024.01 | 86.8 | |
| EDACImplementation=Ours2021.10 | 84.6 | |
| EDAC2024.01 | 84.6 | |
| SAC-NImplementation=Ours2021.10 | 84.5 | |
| SAC-Min2024.01 | 84.5 | |
| CQL-Min2024.01 | 82.07 | |
| BRACImplementation=Reproduced2021.10 | 78 | |
| CQLImplementation=Reproduced2021.10 | 76.9 | |
| MORELImplementation=Reproduced2021.10 | 70.1 | |
| BCQImplementation=Reproduced2021.10 | 69.5 | |
| UWACImplementation=Reproduced2021.10 | 65.1 | |
| BCImplementation=Reproduced2021.10 | 62.9 | |
| BC2024.01 | 62.9 | |
| BEARImplementation=Reproduced2021.10 | 60.1 | |
| REMImplementation=Reproduced2021.10 | 27.8 |