Reinforcement Learning on Atari 2600 FREEWAY
32.4ScoreDQN
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DQNTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=5, Evaluation protocol=Machado et al. (2018a)2018.07 | 32.4 | 0.3 | |
| Average HumanPure-exploration regime=true2022.06 | 29.6 | — | |
| DQN MMCTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=52018.07 | 29.5 | 0.1 | |
| DQN MMC PIXELCNNTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=52018.07 | 29.4 | — | |
| DQN MMC + SRTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=5, w_SR=1000, beta=0.052018.07 | 29.4 | 0.1 | |
| DQN MMC CTSTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=52018.07 | 29.2 | — | |
| ICMPure-exploration regime=true, Episodes=10, Seeds=32022.06 | 15.74 | — | |
| BYOL-ExplorePure-exploration regime=true, Episodes=10, Seeds=32022.06 | 12.94 | — | |
| RNDPure-exploration regime=true, Episodes=10, Seeds=32022.06 | 9.89 | — |