Offline Reinforcement Learning on Antmaze large-diverse v2
94D4RL ScoreInverter
Evaluation Results
| Method | Links | |
|---|---|---|
| Inverterx slower per env step=1.0x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 94 | |
| ReBRACx slower per env step=1.6x†, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=JAX+JIT2026.05 | 64 | |
| IQLx slower per env step=12.6x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 30.2 | |
| CQLx slower per env step=13.7x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 20.5 | |
| BC-10%x slower per env step=9.3x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0.8 | |
| BCx slower per env step=12.7x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| TD3+BCx slower per env step=12.5x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| AWACx slower per env step=9.6x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| SAC-Nx slower per env step=13.4x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| EDACx slower per env step=13.0x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| DTx slower per env step=26.6x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 |