Offline Reinforcement Learning on AntMaze Medium-Play v0
8,830Avg Normalized ScoresfBC
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| sfBCseeds=10, checkpoint_selection=last 10 average2023.10 | 8,830 | — | |
| SEEMseeds=102023.10 | 8,560 | 92 | |
| MSGseeds=10, checkpoint_selection=last 10 average2023.10 | 8,300 | — | |
| IQLseeds=10, checkpoint_selection=last 10 average2023.10 | 7,120 | — | |
| QDQrandom seeds=52024.10 | 81.5 | — | |
| IQLrandom seeds=52024.10 | 71.2 | — | |
| CQLrandom seeds=52024.10 | 61.2 | — | |
| TD3+BCseeds=10, checkpoint_selection=last 10 average2023.10 | 20 | — | |
| TD3+BCrandom seeds=52024.10 | 10.6 | — | |
| Onestep RLrandom seeds=52024.10 | 0.3 | — | |
| diff-QLseeds=102023.10 | 0 | 79.8 | |
| BCrandom seeds=52024.10 | 0 | — | |
| DTrandom seeds=52024.10 | 0 | — | |
| AWACrandom seeds=52024.10 | 0 | — |