Online Fine-Tuning Reinforcement Learning on D4RL Locomotion v2
97.6Hopper v2 MR ScoreIQL + FamO2O
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| IQL + FamO2Omode=Ours (augmented), seeds=62023.10 | 97.6 | 90.7 | 87.3 | 53.1 | 59.2 | 93.1 | 92.9 | 85.5 | 112.7 | 772 | |
| Average (AWAC, IQL) + FamO2Omode=Ours (augmented), seeds=62023.10 | 92.2 | 82.8 | 90.1 | 51 | 53.4 | 91.8 | 88.6 | 82.8 | 110.6 | 743.4 | |
| IQLmode=Base, seeds=62023.10 | 91 | 65.4 | 76.5 | 53.7 | 52.5 | 92.8 | 90.1 | 83.8 | 112.6 | 718.3 | |
| AWAC + FamO2Omode=Ours (augmented), seeds=62023.10 | 86.8 | 75 | 92.9 | 49 | 47.6 | 90.6 | 84.4 | 80 | 108.5 | 714.9 | |
| Average (AWAC, IQL)mode=Base, seeds=62023.10 | 73.5 | 59.7 | 87.1 | 48.8 | 48.7 | 91.9 | 81.5 | 81.4 | 110.9 | 683.4 | |
| AWACmode=Base, seeds=62023.10 | 56 | 54.1 | 97.7 | 43.9 | 44.8 | 91 | 72.8 | 79 | 109.3 | 648.4 |