Offline Reinforcement Learning on D4RL Walker2d Medium
87.7Normalized Avg ReturnMOBILE
Evaluation Results
| Method | Links | |
|---|---|---|
| MOBILEImplementation=Original Paper2024.11 | 87.7 | |
| WFDiffuser2025.09 | 86.1 | |
| MOPOImplementation=Original Paper2024.11 | 84.1 | |
| C-LAPImplementation=Reproduced2024.11 | 82.5 | |
| Decision Diffuser2025.09 | 82.5 | |
| RGG+planning seeds=152023.10 | 82 | |
| RGGplanning seeds=152023.10 | 81.7 | |
| MOBILEImplementation=Reproduced2024.11 | 81 | |
| IQLTraining timesteps=1M, Random seeds=102025.09 | 80.88 | |
| Diffuser2023.10 | 79.7 | |
| PLASImplementation=Reproduced2024.11 | 79.4 | |
| TT2023.10 | 79 | |
| TT2025.09 | 79 | |
| IQL2023.10 | 78.3 | |
| IQL2025.09 | 78.3 | |
| MOPOImplementation=Reproduced2024.11 | 78 | |
| MOREL2023.10 | 77.8 | |
| MOReL2025.09 | 77.8 | |
| BC2023.10 | 75.3 | |
| BC2025.09 | 75.3 | |
| DT2023.10 | 74 | |
| DT2025.09 | 74 | |
| CQL2023.10 | 72.5 | |
| CQL2025.09 | 72.5 | |
| PLASImplementation=Original Paper2024.11 | 44.6 | |
| FQL+BCTraining timesteps=1M, Random seeds=102025.09 | 44.6 | |
| MBOP2023.10 | 41 | |
| MOPO2023.10 | 17.8 |