RL policy training interval optimization on Indoor Lighting Trace Door (simulated 1000 nodes)
89Policy Trainings CountBaseline
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Baseline2018.11 | 89 | 1 | — | |
| Dynamic On-Policy2018.11 | 41 | 0 | 54 |
| Method | Links | |||
|---|---|---|---|---|
| Baseline2018.11 | 89 | 1 | — | |
| Dynamic On-Policy2018.11 | 41 | 0 | 54 |