Continuous Control on BipedalWalker v3
298.4Episodic Cumulative RewardMA-MPPI
Evaluation Results
| Method | Links | |
|---|---|---|
| MA-MPPI2025.09 | 298.4 | |
| AOC-BCsampler=Behavior Clone policy, type=black-box sampler2023.10 | 276.98 | |
| MPPI2025.09 | 241.7 | |
| MPC2025.09 | 219.6 | |
| BCtraining_data=offline dataset of size 1M2023.10 | 208.72 | |
| Data-Avg-Return2023.10 | 202.25 | |
| iLQR2025.09 | 184.2 | |
| SAC2025.09 | 112.6 | |
| PPO2025.09 | 96.3 | |
| DDPG2025.09 | 74.8 | |
| MFRLtype=model-free RL2023.10 | 18.51 | |
| AOC-Uniformsampler=uniform sampler2023.10 | -90.44 | |
| MPCtype=model-based RL2023.10 | -96.82 | |
| KNNvariant=k-neighbors2023.10 | -109.72 | |
| 1NNvariant=nearest-neighbor2023.10 | -111.95 |