Reinforcement Learning on MuJoCo Ant v5
5,953Mean Episodic ReturnSAID
Evaluation Results
| Method | Links | |
|---|---|---|
| SAIDDelay steps=02026.03 | 5,953 | |
| VDPODelay steps=42026.03 | 4,834 | |
| SADRDelay steps=42026.03 | 4,716 | |
| SAIDDelay steps=82026.03 | 4,610 | |
| SAIDDelay steps=42026.03 | 4,586 | |
| SACDelay steps=02026.03 | 4,553 | |
| SADRDelay steps=82026.03 | 4,512 | |
| SAIDDelay steps=162026.03 | 4,137 | |
| SADRDelay steps=162026.03 | 4,118 | |
| VDPODelay steps=82026.03 | 3,032 | |
| VDPODelay steps=162026.03 | 2,588 | |
| DRDelay steps=42026.03 | 2,241 | |
| DRDelay steps=82026.03 | 1,662 | |
| DRDelay steps=162026.03 | 1,282 | |
| SACDelay steps=42026.03 | 1,047 | |
| SACDelay steps=82026.03 | 982 | |
| SACDelay steps=162026.03 | 961 |