Offline Reinforcement Learning on Antmaze umaze v0
98.6Averaged Normalized ScoreMSG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MSGseeds=10, checkpoint_selection=last 10 average2023.10 | 98.6 | — | |
| QDQrandom seeds=52024.10 | 98.6 | — | |
| diff-QLseeds=102023.10 | 95.6 | 96 | |
| SEEMseeds=102023.10 | 94.3 | 97 | |
| sfBCseeds=10, checkpoint_selection=last 10 average2023.10 | 93.3 | — | |
| IQLseeds=10, checkpoint_selection=last 10 average2023.10 | 87.5 | — | |
| IQLrandom seeds=52024.10 | 87.5 | — | |
| TD3+BCrandom seeds=52024.10 | 78.6 | — | |
| CQLrandom seeds=52024.10 | 74 | — | |
| Onestep RLrandom seeds=52024.10 | 64.3 | — | |
| DTrandom seeds=52024.10 | 59.2 | — | |
| AWACrandom seeds=52024.10 | 56.7 | — | |
| BCrandom seeds=52024.10 | 54.6 | — | |
| TD3+BCseeds=10, checkpoint_selection=last 10 average2023.10 | 40.2 | — |