Policy Customization on Humanoid (MuJoCo) (test)
9,306.88Total RewardGreedy
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GreedyLearning Type=IL, Episodes=2002023.06 | 9,306.88 | 6,586.47 | 2.72 | 2,720.41 | |
| Residual-QLearning Type=IL, Episodes=2002023.06 | 7,610.38 | 5,206.52 | 2.41 | 2,403.87 | |
| GreedyLearning Type=RL, Episodes=2002023.06 | 6,209.65 | 4,513.35 | 1.74 | 1,696.3 | |
| Residual-QLearning Type=RL, Episodes=2002023.06 | 6,126.81 | 5,363.79 | 0.76 | 763.02 | |
| Full PolicyLearning Type=RL, Episodes=2002023.06 | 5,771.79 | 5,403.25 | 0.37 | 368.55 | |
| Prior PolicyLearning Type=RL, Episodes=2002023.06 | 5,514.65 | 5,472.68 | 0.04 | 41.97 | |
| Prior PolicyLearning Type=IL, Episodes=200, Note=Trained by BC2023.06 | 4,848.65 | 4,874.01 | -0.01 | -25.35 |