Offline Reinforcement Learning on D4RL HalfCheetah Med-Replay
72.1Normalized Avg ReturnMOPO
Evaluation Results
| Method | Links | |
|---|---|---|
| MOPOImplementation=Original Paper2024.11 | 72.1 | |
| MOBILEImplementation=Original Paper2024.11 | 71.7 | |
| MOPOImplementation=Reproduced2024.11 | 66.3 | |
| MOBILEImplementation=Reproduced2024.11 | 63.7 | |
| C-LAPImplementation=Reproduced2024.11 | 55.5 | |
| MOPO2023.10 | 53.1 | |
| FQL+BCTraining timesteps=1M, Random seeds=102025.09 | 46.23 | |
| CQL2023.10 | 45.5 | |
| CQL2025.09 | 45.5 | |
| IQL2023.10 | 44.2 | |
| IQL2025.09 | 44.2 | |
| PLASImplementation=Original Paper2024.11 | 43.9 | |
| IQLTraining timesteps=1M, Random seeds=102025.09 | 43.44 | |
| MBOP2023.10 | 42.3 | |
| Diffuser2023.10 | 42.2 | |
| TT2023.10 | 41.9 | |
| TT2025.09 | 41.9 | |
| PLASImplementation=Reproduced2024.11 | 41.8 | |
| RGG+planning seeds=152023.10 | 41.3 | |
| RGGplanning seeds=152023.10 | 41 | |
| MOREL2023.10 | 40.2 | |
| MOReL2025.09 | 40.2 | |
| Decision Diffuser2025.09 | 39.3 | |
| WFDiffuser2025.09 | 38.1 | |
| Aaren2024.05 | 37.91 | |
| BC2023.10 | 36.6 | |
| DT2023.10 | 36.6 | |
| BC2025.09 | 36.6 | |
| DT2025.09 | 36.6 | |
| Transformer2024.05 | 36.57 |