Offline Reinforcement Learning on AntMaze Large Diverse v0
47.5Normalized ScoreIQL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| IQLrandom seeds=52024.10 | 47.5 | — | — | |
| QDQrandom seeds=52024.10 | 31.2 | — | — | |
| CQLrandom seeds=52024.10 | 14.9 | — | — | |
| AWACrandom seeds=52024.10 | 1 | — | — | |
| BCrandom seeds=52024.10 | 0 | — | — | |
| TD3+BCrandom seeds=52024.10 | 0 | — | — | |
| DTrandom seeds=52024.10 | 0 | — | — | |
| Onestep RLrandom seeds=52024.10 | 0 | — | — | |
| diff-QLseeds=102023.10 | — | 4.4 | 61.7 | |
| IQLseeds=10, checkpoint_selection=last 10 average2023.10 | — | 47.5 | — | |
| MSGseeds=10, checkpoint_selection=last 10 average2023.10 | — | 58.2 | — | |
| SEEMseeds=102023.10 | — | 67.1 | 75.7 | |
| sfBCseeds=10, checkpoint_selection=last 10 average2023.10 | — | 41.7 | — | |
| TD3+BCseeds=10, checkpoint_selection=last 10 average2023.10 | — | 0 | — |