Offline Reinforcement Learning on D4RL HalfCheetah v2 (random)
1,922.07Average True ReturnSSR-RRS
Evaluation Results
| Method | Links | |
|---|---|---|
| SSR-RRSsplits=2, re-trained on full dataset=true2022.10 | 1,922.07 | |
| SSR-RRSsplits=5, re-trained on full dataset=true2022.10 | 1,922.07 | |
| Optimal Policy2022.10 | 1,922.07 | |
| BVFTvariant=pi + FQE, re-trained on full dataset=true2022.10 | 1,106.94 | |
| CVsplits=2, re-trained on full dataset=true2022.10 | -1.13 | |
| CVsplits=5, re-trained on full dataset=true2022.10 | -1.13 | |
| BVFTvariant=pi x FQE, re-trained on full dataset=true2022.10 | -1.14 |