Policy Customization on Hopper (MuJoCo) (test)
4,798.23Total RewardResidual-Q
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Residual-QLearning Type=RL, Episodes=2002023.06 | 4,798.23 | 3,428.7 | 1.37 | 1,369.52 | |
| Residual-QLearning Type=IL, Episodes=2002023.06 | 4,704.97 | 3,335.57 | 1.37 | 1,369.4 | |
| Full PolicyLearning Type=RL, Episodes=2002023.06 | 4,698.78 | 3,242.15 | 1.46 | 1,456.63 | |
| GreedyLearning Type=RL, Episodes=2002023.06 | 4,661.77 | 3,266.62 | 1.4 | 1,395.15 | |
| GreedyLearning Type=IL, Episodes=2002023.06 | 4,619.94 | 3,236.77 | 1.38 | 1,383.17 | |
| Prior PolicyLearning Type=RL, Episodes=2002023.06 | 4,439.13 | 3,217.66 | 1.33 | 1,221.47 | |
| Prior PolicyLearning Type=IL, Episodes=2002023.06 | 3,828.48 | 2,754.69 | 1.32 | 1,073.79 |