Offline Reinforcement Learning on D4RL Gym hopper-expert
112.8Normalized Avg ReturnRORL
Evaluation Results
| Method | Links | |
|---|---|---|
| RORLImplementation Source=Ours, Ensemble Size=102022.06 | 112.8 | |
| SPQRBase algorithm=SAC-Min2024.01 | 112 | |
| DOSER2026.05 | 111.6 | |
| DMG2026.05 | 111.5 | |
| PDiTSeeds=52026.06 | 111.4 | |
| MORELImplementation=Reproduced2021.10 | 111 | |
| CQLSeeds=52026.06 | 111 | |
| BC2026.05 | 110.9 | |
| PBRL2022.06 | 110.5 | |
| SAC-NImplementation=Ours2021.10 | 110.3 | |
| SAC-Min2024.01 | 110.3 | |
| EDACImplementation=Ours2021.10 | 110.1 | |
| EDACImplementation Source=Paper2022.06 | 110.1 | |
| EDAC2024.01 | 110.1 | |
| CQLImplementation=Original Paper2021.10 | 109.9 | |
| CQL-Min2024.01 | 109.9 | |
| DTSeeds=52026.06 | 109.8 | |
| AWAC2026.05 | 109.5 | |
| IQL2026.05 | 109.4 | |
| BCQImplementation=Reproduced2021.10 | 109 | |
| BCQ2026.05 | 109 | |
| TD3+BC2026.05 | 107.8 | |
| BCImplementation=Reproduced2021.10 | 107.7 | |
| BC2022.06 | 107.7 | |
| BC2024.01 | 107.7 | |
| OneStep2026.05 | 106.9 | |
| CQLImplementation=Reproduced2021.10 | 106.5 | |
| CQL2022.06 | 106.5 | |
| DRIVESeeds=52026.06 | 106.1 | |
| TD3+BCSeeds=52026.06 | 98 | |
| CQL2026.05 | 96.5 | |
| BEARSeeds=52026.06 | 96.3 | |
| DT2026.05 | 94.2 | |
| IQLSeeds=52026.06 | 91.5 | |
| BRACImplementation=Reproduced2021.10 | 78.1 | |
| EDAC-10Implementation Source=Reproduced, Ensemble Size=102022.06 | 77 | |
| BCSeeds=52026.06 | 67.5 | |
| BEAR2026.05 | 54.6 | |
| BEARImplementation=Reproduced2021.10 | 39.4 | |
| UWACImplementation=Reproduced2021.10 | 38.1 | |
| SAC-10Implementation Source=Reproduced, Ensemble Size=102022.06 | 1.1 | |
| REMImplementation=Reproduced2021.10 | 0.8 | |
| SACImplementation=Reproduced2021.10 | 0.7 |