Reinforcement Learning on MuJoCo Hopper
3,876Average ReturnSAC
Evaluation Results
| Method | Links | |
|---|---|---|
| SAC2023.05 | 3,876 | |
| SPMD2023.05 | 3,619 | |
| TD3Training Steps=1M2026.03 | 3,387 | |
| Oracle-TC M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 3,281.92 | |
| Stacked TC M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 3,124.06 | |
| Oracle-TC RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 3,071.17 | |
| SiMPO-ExpTraining Steps=1M2026.03 | 3,002 | |
| SiMPO-SquareTraining Steps=1M2026.03 | 2,809 | |
| Vanilla-TC M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 2,756.5 | |
| SiMPO-LinearTraining Steps=1M2026.03 | 2,637 | |
| SiMPO-Lin. Neg.Training Steps=1M2026.03 | 2,609 | |
| SACTraining Steps=1M2026.03 | 2,126 | |
| Stacked TC RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 2,072.75 | |
| DACERTraining Steps=1M2026.03 | 1,813 | |
| DREvaluation Protocol=Raw static average case, Seeds=102024.06 | 1,688.36 | |
| QSMTraining Steps=1M2026.03 | 1,594 | |
| Vanilla-TC RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 1,558.26 | |
| Oracle M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 1,249.62 | |
| M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 1,193.32 | |
| DIPOTraining Steps=1M2026.03 | 1,191 | |
| QVPOTraining Steps=1M2026.03 | 960 | |
| VanillaEvaluation Protocol=Raw static average case, Seeds=102024.06 | 733.18 | |
| Oracle RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 380.39 | |
| RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 276.37 |