RL policy training interval optimization on Indoor Lighting Trace Middle Office (simulated 1000 nodes)
42Policy Trainings CountDynamic On-Policy
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Dynamic On-Policy2018.11 | 42 | 0 | 53 | |
| Baseline2018.11 | 89 | 3 | — |
| Method | Links | |||
|---|---|---|---|---|
| Dynamic On-Policy2018.11 | 42 | 0 | 53 | |
| Baseline2018.11 | 89 | 3 | — |