Offline Reinforcement Learning on D4RL HalfCheetah Med-Replay v2
52.2Avg Normalized ReturnSPOT
Evaluation Results
| Method | Links | |
|---|---|---|
| SPOT2022.02 | 52.2 | |
| BRAC2021.06 | 47.7 | |
| Fu et al.Algorithmic Template=Iterative2021.06 | 47.7 | |
| CQL2021.06 | 45.5 | |
| CQL2022.02 | 45.5 | |
| TD3+BC2022.02 | 44.6 | |
| IQL2022.02 | 44.2 | |
| Trajectory Transformerdiscretization=uniform2021.06 | 44.1 | |
| Rev. KL Reg.Algorithmic Template=One-step2021.06 | 42.4 | |
| MBOP2021.06 | 42.3 | |
| Trajectory Transformerdiscretization=quantile2021.06 | 41.9 | |
| AWAC2022.02 | 40.5 | |
| GACStrategy=Exploitation query (GAC-E[y])2025.12 | 39.8 | |
| LPT2025.12 | 39.6 | |
| GACStrategy=Exploration query (GAC-p(y+))2025.12 | 38.8 | |
| Exp. WeightAlgorithmic Template=One-step2021.06 | 38.6 | |
| Easy BCQAlgorithmic Template=One-step2021.06 | 38.4 | |
| Onestep2022.02 | 38.1 | |
| GACStrategy=Fixed target steering (GAC-y*)2025.12 | 36.7 | |
| Decision Transformer2021.06 | 36.6 | |
| BC2022.02 | 36.6 | |
| DT2022.02 | 36.6 | |
| BCAlgorithmic Template=One-step2021.06 | 34.9 | |
| DT2025.12 | 33.3 | |
| GACStrategy=Sampling from prior (p(y|z)p(z))2025.12 | 33 | |
| QDT2025.12 | 32.8 | |
| CQL2025.12 | 7.8 | |
| IQL2025.12 | 5.2 | |
| Behavior Cloning2021.06 | 4.3 |