Reinforcement Learning on MuJoCo Ant
7,889.1Average ReturnOracle-TC RARL
Evaluation Results
| Method | Links | |
|---|---|---|
| Oracle-TC RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 7,889.1 | |
| Oracle-TC M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 7,739.65 | |
| Vanilla-TC RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 7,558.58 | |
| DREvaluation Protocol=Raw static average case, Seeds=102024.06 | 7,500.88 | |
| Vanilla-TC M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 7,366.9 | |
| Stacked TC RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 7,123.07 | |
| Stacked TC M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 6,912.76 | |
| SiMPO-LinearTraining Steps=1M2026.03 | 5,984 | |
| Oracle M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 5,958.21 | |
| SiMPO-SquareTraining Steps=1M2026.03 | 5,925 | |
| SiMPO-Lin. Neg.Training Steps=1M2026.03 | 5,700 | |
| M2TD3Evaluation Protocol=Raw static average case, Seeds=102024.06 | 5,577.41 | |
| SPMD2023.05 | 5,230 | |
| SAC2023.05 | 5,118 | |
| SiMPO-ExpTraining Steps=1M2026.03 | 5,071 | |
| DACERTraining Steps=1M2026.03 | 4,887 | |
| Oracle RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 4,684.83 | |
| RARLEvaluation Protocol=Raw static average case, Seeds=102024.06 | 4,650.55 | |
| TD3Training Steps=1M2026.03 | 4,400 | |
| SACTraining Steps=1M2026.03 | 2,858 | |
| VanillaEvaluation Protocol=Raw static average case, Seeds=102024.06 | 2,600.43 | |
| QVPOTraining Steps=1M2026.03 | 2,122 | |
| DIPOTraining Steps=1M2026.03 | 977 | |
| QSMTraining Steps=1M2026.03 | 711 |