Offline Reinforcement Learning on AntMaze Medium-Diverse v2
6.6Average ScoreTRACER
Evaluation Results
| Method | Links | |
|---|---|---|
| TRACERcorruption_type=random simultaneous corruptions2024.11 | 6.6 | |
| Inverterx slower per env step=1.0x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0.965 | |
| ReBRACx slower per env step=1.6x†, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=JAX+JIT2026.05 | 0.835 | |
| IQLx slower per env step=12.9x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0.738 | |
| CQLx slower per env step=14.0x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0.672 | |
| BC-10%x slower per env step=9.5x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0.058 | |
| BCx slower per env step=13.0x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0.008 | |
| TD3+BCx slower per env step=12.7x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0.002 | |
| IQLcorruption_type=random simultaneous corruptions2024.11 | 0 | |
| RIQLcorruption_type=random simultaneous corruptions2024.11 | 0 | |
| AWACx slower per env step=9.8x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| SAC-Nx slower per env step=13.7x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| EDACx slower per env step=13.3x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| DTx slower per env step=27.3x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 |