Reinforcement Learning on busyRingMC4
6.08TimeMungojerrie
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Mungojerriestates=2592, prod.=15426, c=-1, ε=0.1, α=0.1, η=0.01, train-steps=1.5M2025.05 | 6.08 | — | — | |
| Q-learning with reduction (Bozkurt et al. 2020)states=2592, prod.=15426, c=-1, ε=0.1, α=0.1, η=0.01, train-steps=1.5M2025.05 | — | — | 2.33 | |
| Q-learning with reduction (Hahn et al. 2019)states=2592, prod.=15426, c=-1, ε=0.1, α=0.1, η=0.01, train-steps=1.5M2025.05 | — | 3.94 | — |