Offline Reinforcement Learning on AntMaze Medium-Play v2
89.5Average ScoreReBRAC
Evaluation Results
| Method | Links | |
|---|---|---|
| ReBRACx slower per env step=1.6x†, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=JAX+JIT2026.05 | 89.5 | |
| Inverterx slower per env step=1.0x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 87.8 | |
| CQLx slower per env step=13.8x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 65.8 | |
| IQLx slower per env step=12.7x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 65.8 | |
| TRACERcorruption_type=random simultaneous corruptions2024.11 | 7.5 | |
| BC-10%x slower per env step=9.4x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 2 | |
| TD3+BCx slower per env step=12.6x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0.2 | |
| IQLcorruption_type=random simultaneous corruptions2024.11 | 0 | |
| RIQLcorruption_type=random simultaneous corruptions2024.11 | 0 | |
| BCx slower per env step=12.8x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| AWACx slower per env step=9.6x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| SAC-Nx slower per env step=13.5x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| EDACx slower per env step=13.1x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 | |
| DTx slower per env step=26.8x, Hardware=single A40 GPU, Batch size=1, Number of episodes=100, Implementation framework=PyTorch2026.05 | 0 |