Offline Reinforcement Learning on Antmaze umaze-diverse v0
88.5Avg Normalized ScoreSEEM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SEEMseeds=102023.10 | 88.5 | 95 | |
| sfBCseeds=10, checkpoint_selection=last 10 average2023.10 | 86.7 | — | |
| CQLrandom seeds=52024.10 | 84 | — | |
| MSGseeds=10, checkpoint_selection=last 10 average2023.10 | 76.7 | — | |
| TD3+BCrandom seeds=52024.10 | 71.4 | — | |
| diff-QLseeds=102023.10 | 69.5 | 84 | |
| QDQrandom seeds=52024.10 | 67.8 | — | |
| IQLseeds=10, checkpoint_selection=last 10 average2023.10 | 62.2 | — | |
| IQLrandom seeds=52024.10 | 62.2 | — | |
| Onestep RLrandom seeds=52024.10 | 60.7 | — | |
| TD3+BCseeds=10, checkpoint_selection=last 10 average2023.10 | 58 | — | |
| DTrandom seeds=52024.10 | 53 | — | |
| AWACrandom seeds=52024.10 | 49.3 | — | |
| BCrandom seeds=52024.10 | 45.6 | — |