Reinforcement Learning on MuJoCo Half-Cheetah
13,907Average ReturnSiMPO-Lin. Neg.
Evaluation Results
| Method | Links | |
|---|---|---|
| SiMPO-Lin. Neg.Training Steps=1M2026.03 | 13,907 | |
| SAC2023.05 | 13,300 | |
| SPMD2023.05 | 13,025 | |
| TD3Training Steps=1M2026.03 | 9,820 | |
| Oracle-TC M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 9,536.92 | |
| Oracle-TC RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 9,474 | |
| Stacked TC M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 8,583.55 | |
| Optimal in TargetMode=Optimal Oracle2024.11 | 8,543 | |
| Vanilla-TC M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 8,467.64 | |
| QVPOTraining Steps=1M2026.03 | 8,081 | |
| DARAILMode=Evaluation in Target2024.11 | 7,067 | |
| DARCMode=Training in Source2024.11 | 6,995 | |
| DREvaluation Protocol=Raw static average case, Seeds=102024.06 | 6,170.33 | |
| Stacked TC RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 6,130.71 | |
| Vanilla-TC RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 6,092.61 | |
| Oracle M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 4,930.18 | |
| DARCMode=Evaluation in Target2024.11 | 4,133 | |
| M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 4,000.98 | |
| VanillaEvaluation Protocol=Raw static average case, Seeds=102024.06 | 2,350.58 | |
| RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 206.71 | |
| Oracle RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 36.19 | |
| DACERTraining Steps=1M2026.03 | 13 | |
| SiMPO-ExpTraining Steps=1M2026.03 | 13 | |
| SiMPO-SquareTraining Steps=1M2026.03 | 13 | |
| SiMPO-LinearTraining Steps=1M2026.03 | 13 | |
| QSMTraining Steps=1M2026.03 | 12 | |
| SACTraining Steps=1M2026.03 | 10 | |
| DIPOTraining Steps=1M2026.03 | 10 |