Offline Reinforcement Learning on D4RL Walker2D Expert
117.4Mean Normalized ScoreAnti-exploration Method
Evaluation Results
| Method | Links | |
|---|---|---|
| Anti-exploration Method2026.02 | 117.4 | |
| PMDBNumber of random seeds=42026.03 | 115.9 | |
| PMDB2026.05 | 115.9 | |
| DMG2026.02 | 114.7 | |
| DMGNumber of random seeds=42026.03 | 114.7 | |
| DMG2026.05 | 114.7 | |
| SAC-RND2026.02 | 114.3 | |
| MOPOImplementation=Reproduced2024.11 | 113.4 | |
| C-LAPImplementation=Reproduced2024.11 | 111.7 | |
| GPC-SAC2026.02 | 111.7 | |
| RRPINumber of random seeds=42026.03 | 111.2 | |
| IQL2026.02 | 109.9 | |
| EPQNumber of random seeds=42026.03 | 109.8 | |
| EPQ2026.05 | 109.8 | |
| PSPOAveraged across=4 random seeds2026.05 | 109.8 | |
| CQLNumber of random seeds=42026.03 | 109.3 | |
| CQL2026.05 | 109.3 | |
| PLASImplementation=Reproduced2024.11 | 109.1 | |
| BooT-rGeneration Strategy=Autoregressive2022.06 | 108.7 | |
| BooT-oGeneration Strategy=Autoregressive2022.06 | 108.5 | |
| BooT-oGeneration Strategy=Teacher-forcing2022.06 | 108.5 | |
| BooT-rGeneration Strategy=Teacher-forcing2022.06 | 108.5 | |
| CQL2026.02 | 108.5 | |
| TT2022.06 | 108.4 | |
| IGDF2025.12 | 108.3 | |
| DROCO2025.12 | 104.5 | |
| OTDF2025.12 | 103.5 | |
| IQL*variant=tuned/reproduced2025.12 | 103.4 | |
| DARA2025.12 | 102.7 | |
| CQL*variant=tuned/reproduced2025.12 | 79 | |
| MOReLNumber of random seeds=42026.03 | 62.6 | |
| MOReL2026.05 | 62.6 | |
| BOSA2025.12 | 30.2 | |
| MOBILEImplementation=Reproduced2024.11 | 15 | |
| AMGNumber of random seeds=42026.03 | 5.5 | |
| ADM2026.05 | 5.5 | |
| RAMBONumber of random seeds=42026.03 | 1.6 | |
| RAMBO2026.05 | 1.6 |