Offline Reinforcement Learning on D4RL Gym halfcheetah-expert
112.8Normalized ReturnSPQR
Evaluation Results
| Method | Links | |
|---|---|---|
| SPQRBase algorithm=SAC-Min2024.01 | 112.8 | |
| SPQRBase algorithm=EDAC2024.01 | 110.3 | |
| EDACImplementation=Ours2021.10 | 106.8 | |
| EDACImplementation Source=Paper2022.06 | 106.8 | |
| EDAC2024.01 | 106.8 | |
| SAC-NImplementation=Ours2021.10 | 105.2 | |
| RORLImplementation Source=Ours, Ensemble Size=102022.06 | 105.2 | |
| SAC-Min2024.01 | 105.2 | |
| SAC-10Implementation Source=Reproduced, Ensemble Size=102022.06 | 104.9 | |
| CQLImplementation=Original Paper2021.10 | 104.8 | |
| CQL-Min2024.01 | 104.8 | |
| EDAC-10Implementation Source=Reproduced, Ensemble Size=102022.06 | 104 | |
| CQLImplementation=Reproduced2021.10 | 97.3 | |
| CQL2022.06 | 97.3 | |
| TD3+BC2026.05 | 96.7 | |
| CQL2026.05 | 96.3 | |
| DMG2026.05 | 95.9 | |
| DOSER2026.05 | 95.4 | |
| IQL2026.05 | 95 | |
| UWACImplementation=Reproduced2021.10 | 94.7 | |
| BC2026.05 | 92.9 | |
| BEAR2026.05 | 92.7 | |
| BEARImplementation=Reproduced2021.10 | 92.6 | |
| PBRL2022.06 | 92.4 | |
| BCImplementation=Reproduced2021.10 | 91.8 | |
| BC2022.06 | 91.8 | |
| BC2024.01 | 91.8 | |
| BCQImplementation=Reproduced2021.10 | 89.9 | |
| BCQ2026.05 | 89.9 | |
| OneStep2026.05 | 88.2 | |
| DT2026.05 | 87.7 | |
| AWAC2026.05 | 81.7 | |
| BRACImplementation=Reproduced2021.10 | 39 | |
| MORELImplementation=Reproduced2021.10 | 8.4 | |
| REMImplementation=Reproduced2021.10 | 4.1 | |
| SACImplementation=Reproduced2021.10 | -0.8 |