Continuous Locomotion Control on MuJoCo Ant v4 (train)
2.43Normalized Final ReturnDUPO
Evaluation Results
| Method | Links | |
|---|---|---|
| DUPOMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 2.43 | |
| State PredictionMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 2.28 | |
| DUPOMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 2.14 | |
| State PredictionMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 2.11 | |
| State PredictionMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1.9 | |
| DUPOMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1.87 | |
| State AugmentationMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 1.57 | |
| State AugmentationMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1.43 | |
| State AugmentationMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1.3 | |
| DC/ACMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1 | |
| DC/ACMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1 | |
| DC/ACMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 1 |