Offline Reinforcement Learning on AntMaze large-diverse (l-d)
81.8Normalized ScoreVIPO-LEQ
Evaluation Results
| Method | Links | |
|---|---|---|
| VIPO-LEQ2025.04 | 81.8 | |
| QCS-GAlgorithm Category=Ours, Conditioning Strategy=Goal2024.02 | 77.3 | |
| PORAlgorithm Category=Combined2024.02 | 73.4 | |
| QCS-RAlgorithm Category=Ours, Conditioning Strategy=Return2024.02 | 70.6 | |
| LEQ2025.04 | 60.2 | |
| SQLAlgorithm Category=Value-Based Method2024.02 | 52.3 | |
| IQLAlgorithm Category=Value-Based Method2024.02 | 47.5 | |
| RvS-GAlgorithm Category=RCSL, Conditioning Strategy=Goal2024.02 | 36.9 | |
| CQLAlgorithm Category=Value-Based Method2024.02 | 14.9 | |
| DCAlgorithm Category=RCSL2024.02 | 12.3 | |
| IQL-TD-MPC2025.04 | 4 | |
| RvS-RAlgorithm Category=RCSL, Conditioning Strategy=Return2024.02 | 3.7 | |
| DTAlgorithm Category=RCSL2024.02 | 0.5 | |
| TD3+BCAlgorithm Category=Value-Based Method2024.02 | 0 | |
| MOBILE2025.04 | 0 |