Continuous Control on InvertedPendulum v5
1,000Average Episodic RewardTD3
Evaluation Results
| Method | Links | |
|---|---|---|
| TD3NFE=12026.04 | 1,000 | |
| MaxEntDPNFE=20 × 102026.04 | 1,000 | |
| TRFP(ours)NFE=4 × 42026.04 | 1,000 | |
| TRFP(one-step)NFE=1 × 42026.04 | 1,000 | |
| SDACNFE=20 × 322026.04 | 992.1 | |
| PPOEnvironment steps=1M, Seeds=10, Tests per epoch=102026.03 | 975.8 | |
| PDAEnvironment steps=1M, Seeds=10, Tests per epoch=102026.03 | 916 | |
| SACNFE=12026.04 | 808.1 |