Offline Reinforcement Learning on D4RL Hopper (Expert)
116.6Mean Normalized ScorePSPO
Evaluation Results
| Method | Links | |
|---|---|---|
| PSPOAveraged across=4 random seeds2026.05 | 116.6 | |
| RRPINumber of random seeds=42026.03 | 114.8 | |
| OTDF2025.12 | 113.2 | |
| EPQNumber of random seeds=42026.03 | 112.4 | |
| EPQ2026.05 | 112.4 | |
| PMDBNumber of random seeds=42026.03 | 111.7 | |
| PMDB2026.05 | 111.7 | |
| DMGNumber of random seeds=42026.03 | 111.5 | |
| DMG2026.05 | 111.5 | |
| BooT-rGeneration Strategy=Autoregressive2022.06 | 110.5 | |
| C-LAPImplementation=Reproduced2024.11 | 110.5 | |
| BooT-oGeneration Strategy=Autoregressive2022.06 | 110.3 | |
| BooT-rGeneration Strategy=Teacher-forcing2022.06 | 108.2 | |
| CQLNumber of random seeds=42026.03 | 106.5 | |
| CQL2026.05 | 106.5 | |
| BooT-oGeneration Strategy=Teacher-forcing2022.06 | 104.6 | |
| TT2022.06 | 102.3 | |
| AMGNumber of random seeds=42026.03 | 102.3 | |
| ADM2026.05 | 102.3 | |
| PLASImplementation=Reproduced2024.11 | 97.1 | |
| DROCO2025.12 | 92.5 | |
| IQL*variant=tuned/reproduced2025.12 | 87.2 | |
| MOReLNumber of random seeds=42026.03 | 80.4 | |
| MOReL2026.05 | 80.4 | |
| MOBILEImplementation=Reproduced2024.11 | 78.6 | |
| DARA2025.12 | 77.1 | |
| CQL*variant=tuned/reproduced2025.12 | 67.9 | |
| BOSA2025.12 | 64.3 | |
| IGDF2025.12 | 51.5 | |
| RAMBONumber of random seeds=42026.03 | 50 | |
| RAMBO2026.05 | 50 | |
| MOPOImplementation=Reproduced2024.11 | 26.2 |