Reinforcement Learning on doublegrid
16.43TimeMungojerrie
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Mungojerriestates=1296, prod.=5183, c=-2, ε=0.5, α=0.05, η=0.01, train-steps=12M2025.05 | 16.43 | — | — | |
| Q-learning with reduction (Bozkurt et al. 2020)states=1296, prod.=5183, c=-2, ε=0.5, α=0.05, η=0.01, train-steps=12M2025.05 | — | — | 3.09 | |
| Q-learning with reduction (Hahn et al. 2019)states=1296, prod.=5183, c=-2, ε=0.5, α=0.05, η=0.01, train-steps=12M2025.05 | — | 3.45 | — |