Offline Inverse Reinforcement Learning on MuJoCo hopper (medium-replay)
3,512.09Average RewardExpert Performance
Evaluation Results
| Method | Links | |
|---|---|---|
| Expert Performanceexpert demonstrations=50002023.02 | 3,512.09 | |
| ValueDICEexpert demonstrations=50002023.02 | 3,073.16 | |
| Offline ML-IRLexpert demonstrations=50002023.02 | 3,046.36 | |
| CLAREexpert demonstrations=50002023.02 | 2,888.04 | |
| BCexpert demonstrations=50002023.02 | 2,801.19 |