Continuous robotic locomotion control on MuJoCo Swimmer v4 (Normalized final return)
2.98Normalized Final ReturnDUPO
Evaluation Results
| Method | Links | |
|---|---|---|
| DUPOMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 2.98 | |
| State AugmentationMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 2.75 | |
| DUPOMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 2.36 | |
| State PredictionMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 2.3 | |
| State AugmentationMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 2.15 | |
| State PredictionMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1.75 | |
| State PredictionMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1.65 | |
| DUPOMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1.25 | |
| State AugmentationMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1.07 | |
| DC/ACMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1 | |
| DC/ACMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1 | |
| DC/ACMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 1 |