Reinforcement Learning on InvManagementBacklogEnv v0
496.4Average Episodic RewardPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| PPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 496.4 | |
| PDAEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 491.6 | |
| NPGEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 430.4 | |
| TRPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 397.5 |