Offline Reinforcement Learning on D4RL Adroit pen human v0
76.3Normalized ReturnDPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| DPPOFramework=Preference Transformer (PT)2023.01 | 76.3 | |
| CQLImplementation Source=Paper2021.10 | 55.8 | |
| PT+IQLFramework=Preference Transformer (PT)2023.01 | 53 | |
| EDACImplementation Source=Ours2021.10 | 52.1 | |
| CQLImplementation Source=Reproduced2021.10 | 35.2 | |
| PT+CQLFramework=Preference Transformer (PT)2023.01 | 31.6 | |
| BC2021.10 | 25.8 | |
| PT+%BCFramework=Preference Transformer (PT)2023.01 | 19.4 | |
| SAC-NImplementation Source=Ours2021.10 | 9.5 | |
| REM2021.10 | 5.4 | |
| SAC2021.10 | 4.3 | |
| PT+RvSFramework=Preference Transformer (PT)2023.01 | -1.8 |