Continuous Control on MuJoCo Reacher
6.22Average RewardTRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| TRPO2026.04 | 6.22 | |
| PPO2026.04 | 5.17 | |
| PolyGRADGuide=None2026.04 | 4.48 | |
| TOP-TD3Training steps=1M, Number of trials=102021.02 | -3.85 | |
| AGD-MBRLGuide=SAG2026.04 | -3.87 | |
| AGD-MBRLGuide=EAG2026.04 | -3.9 | |
| ND TOP-TD3Training steps=1M, Number of trials=102021.02 | -3.91 | |
| QR-TD3Training steps=1M, Number of trials=102021.02 | -3.95 | |
| PolyGRADGuide=Diffuser Guide2026.04 | -3.96 | |
| SACTraining steps=1M, Number of trials=102021.02 | -4.14 | |
| OACTraining steps=1M, Number of trials=102021.02 | -4.15 | |
| TD3Training steps=1M, Number of trials=102021.02 | -4.22 | |
| Implementation [21]strategy=ignore2023.08 | -6.1 | |
| Implementation [20]strategy=ignore2023.08 | -7.2 | |
| Implementation [21]strategy=underest2023.08 | -7.2 | |
| Implementation [20]strategy=underest2023.08 | -10.3 | |
| Implementation [20]strategy=zero2023.08 | -12 | |
| Implementation [21]strategy=zero2023.08 | -15.2 |