Offline Reinforcement Learning on D4RL HalfCheetah (Expert)
108.8Mean Normalized ScoreAnti-exploration Method
Evaluation Results
| Method | Links | |
|---|---|---|
| Anti-exploration Method2026.02 | 108.8 | |
| EPQNumber of random seeds=42026.03 | 107.2 | |
| EPQ2026.05 | 107.2 | |
| SAC-RND2026.02 | 105.8 | |
| PMDBNumber of random seeds=42026.03 | 105.7 | |
| PMDB2026.05 | 105.7 | |
| GPC-SAC2026.02 | 104.7 | |
| MOBILEImplementation=Reproduced2024.11 | 103 | |
| CQLNumber of random seeds=42026.03 | 97.3 | |
| CQL2026.05 | 97.3 | |
| C-LAPImplementation=Reproduced2024.11 | 97.1 | |
| CQL2026.02 | 96.3 | |
| DMG2026.02 | 95.9 | |
| DMGNumber of random seeds=42026.03 | 95.9 | |
| DMG2026.05 | 95.9 | |
| BooT-rGeneration Strategy=Teacher-forcing2022.06 | 95.4 | |
| TT2022.06 | 95.3 | |
| BooT-oGeneration Strategy=Teacher-forcing2022.06 | 95 | |
| IQL2026.02 | 95 | |
| PLASImplementation=Reproduced2024.11 | 94.8 | |
| BooT-rGeneration Strategy=Autoregressive2022.06 | 94.4 | |
| MOPOImplementation=Reproduced2024.11 | 94 | |
| PSPOAveraged across=4 random seeds2026.05 | 93.3 | |
| BooT-oGeneration Strategy=Autoregressive2022.06 | 92.3 | |
| RRPINumber of random seeds=42026.03 | 90.7 | |
| AMGNumber of random seeds=42026.03 | 89.4 | |
| ADM2026.05 | 89.4 | |
| RAMBONumber of random seeds=42026.03 | 79.3 | |
| RAMBO2026.05 | 79.3 | |
| MOReLNumber of random seeds=42026.03 | 8.4 | |
| MOReL2026.05 | 8.4 |