Continuous Control on Swimmer v5
119.3Average Episodic RewardTRFP(ours)
Evaluation Results
| Method | Links | |
|---|---|---|
| TRFP(ours)NFE=4 × 42026.04 | 119.3 | |
| TRFP(one-step)NFE=1 × 42026.04 | 118.9 | |
| PDAEnvironment steps=1M, Seeds=10, Tests per epoch=102026.03 | 111.3 | |
| SDACNFE=20 × 322026.04 | 85.8 | |
| MaxEntDPNFE=20 × 102026.04 | 75.8 | |
| TD3NFE=12026.04 | 58.9 | |
| SACNFE=12026.04 | 58.5 | |
| PPOEnvironment steps=1M, Seeds=10, Tests per epoch=102026.03 | 54.2 |