Offline Reinforcement Learning on D4RL Gym halfcheetah-medium
74.8Normalized ReturnSPQR
Evaluation Results
| Method | Links | |
|---|---|---|
| SPQRBase algorithm=SAC-Min2024.01 | 74.8 | |
| MOPOalgorithm class=model-based2021.10 | 69.5 | |
| SAC-NImplementation=Ours2021.10 | 67.5 | |
| SAC-Min2024.01 | 67.5 | |
| RORLImplementation Source=Ours, Ensemble Size=102022.06 | 66.8 | |
| EDACImplementation=Ours2021.10 | 65.9 | |
| EDACImplementation Source=Paper2022.06 | 65.9 | |
| EDAC2024.01 | 65.9 | |
| SAC-10Implementation Source=Reproduced, Ensemble Size=102022.06 | 64.9 | |
| EDAC-10Implementation Source=Reproduced, Ensemble Size=102022.06 | 64.1 | |
| MORELImplementation=Reproduced2021.10 | 60.7 | |
| PBRL2022.06 | 57.9 | |
| FQL2025.10 | 55.6 | |
| SACImplementation=Reproduced2021.10 | 55.2 | |
| QIPO-OT2025.10 | 54.2 | |
| GTP2025.10 | 53.9 | |
| SSCQL2025.10 | 52.3 | |
| BRACImplementation=Reproduced2021.10 | 51.9 | |
| IPDEvaluation Protocol=Tenfold Episode Evaluation2026.03 | 51.2 | |
| TTConfiguration=Trifle2023.10 | 49.5 | |
| QQLDomain=Gym Locomotion, Hyperparameter Tuning=consistent2025.11 | 49.5 | |
| CQLEvaluation Protocol=Tenfold Episode Evaluation2026.03 | 49.2 | |
| ROMI-CQLbase algorithm=CQL2021.10 | 49.1 | |
| Decision Diffuser2024.06 | 49.1 | |
| DD2023.10 | 49.1 | |
| DDEvaluation Protocol=Tenfold Episode Evaluation2026.03 | 49.1 | |
| QTEvaluation Protocol=Tenfold Episode Evaluation2026.03 | 49.1 | |
| CQLalgorithm class=model-free2021.10 | 49 | |
| TT(+Q)Configuration=Trifle2023.10 | 48.9 | |
| TT(+Q)Configuration=base2023.10 | 48.7 | |
| TD3(+BC)2023.10 | 48.3 | |
| TD3+BCDomain=Gym Locomotion, Hyperparameter Tuning=individually tuned2025.11 | 48.3 | |
| XQLDomain=Gym Locomotion, Hyperparameter Tuning=individually tuned2025.11 | 48.3 | |
| QIPO-Diff2025.10 | 48.2 | |
| MXQLDomain=Gym Locomotion, Hyperparameter Tuning=individually tuned2025.11 | 47.7 | |
| IQL2023.10 | 47.4 | |
| IQLEvaluation Protocol=Tenfold Episode Evaluation2026.03 | 47.4 | |
| IQLDomain=Gym Locomotion, Hyperparameter Tuning=individually tuned2025.11 | 47.4 | |
| CQLImplementation=Reproduced2021.10 | 46.9 | |
| CQL2022.06 | 46.9 | |
| Trajectory Transformer2024.06 | 46.9 | |
| TTConfiguration=base2023.10 | 46.9 | |
| BCQImplementation=Reproduced2021.10 | 46.6 | |
| CQLImplementation=Original Paper2021.10 | 44.4 | |
| CQL-Min2024.01 | 44.4 | |
| Diffuser2024.06 | 44.2 | |
| DTConfiguration=Trifle2023.10 | 44.2 | |
| CQL2023.10 | 44 | |
| CQLDomain=Gym Locomotion, Hyperparameter Tuning=individually tuned2025.11 | 44 | |
| Decision Mamba2024.06 | 43.8 | |
| UWACImplementation=Reproduced2021.10 | 43.7 | |
| BCImplementation=Reproduced2021.10 | 43.2 | |
| BC2022.06 | 43.2 | |
| BC2024.01 | 43.2 | |
| Critic-Guided Decision Transformer2024.06 | 43 | |
| ReinformerEvaluation Protocol=Tenfold Episode Evaluation2026.03 | 42.9 | |
| BEARImplementation=Reproduced2021.10 | 42.8 | |
| DTConfiguration=base2023.10 | 42.6 | |
| DTEvaluation Protocol=Tenfold Episode Evaluation2026.03 | 42.6 | |
| BCDomain=Gym Locomotion, Hyperparameter Tuning=individually tuned2025.11 | 42.6 | |
| %BC2023.10 | 42.5 | |
| EDTEvaluation Protocol=Tenfold Episode Evaluation2026.03 | 42.5 | |
| QDTEvaluation Protocol=Tenfold Episode Evaluation2026.03 | 39.3 | |
| BC2021.10 | 39.2 | |
| REMImplementation=Reproduced2021.10 | -0.8 |