Continuous Control on MuJoCo HalfCheetah v4
107Normalized PerformanceAD-SAC
Evaluation Results
| Method | Links | |
|---|---|---|
| AD-SACDelays=5, Global steps=1M2024.05 | 107 | |
| VDPODelays=5, Global steps=1M2024.05 | 103 | |
| BPQLDelays=5, Global steps=1M2024.05 | 100 | |
| DIDADelays=5, Global steps=1M2024.05 | 90 | |
| BPQLDelays=25, Global steps=1M2024.05 | 87 | |
| AD-SACDelays=50, Global steps=1M2024.05 | 74 | |
| BPQLDelays=50, Global steps=1M2024.05 | 73 | |
| VDPODelays=50, Global steps=1M2024.05 | 72 | |
| AD-SACDelays=25, Global steps=1M2024.05 | 71 | |
| VDPODelays=25, Global steps=1M2024.05 | 70 | |
| DC/ACDelays=5, Global steps=1M2024.05 | 40 | |
| A-SACDelays=5, Global steps=1M2024.05 | 35 | |
| DC/ACDelays=25, Global steps=1M2024.05 | 16 | |
| DUPOMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 15.96 | |
| DIDADelays=50, Global steps=1M2024.05 | 15 | |
| State PredictionMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 13.19 | |
| DIDADelays=25, Global steps=1M2024.05 | 12 | |
| A-SACDelays=50, Global steps=1M2024.05 | 12 | |
| DC/ACDelays=50, Global steps=1M2024.05 | 12 | |
| DUPOMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 7.53 | |
| State PredictionMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 6.42 | |
| State AugmentationMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 5.96 | |
| State AugmentationMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 5.1 | |
| DUPOMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 4.53 | |
| A-SACDelays=25, Global steps=1M2024.05 | 4 | |
| State PredictionMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 3.93 | |
| State AugmentationMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 3.78 | |
| DC/ACMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1 | |
| DC/ACMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1 | |
| DC/ACMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 1 |