Reinforcement Learning on Atari 2600 57 games (test)
1,006.4Median Human-Normalized ScoreMuZero Res2 Adam
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| MuZero Res2 AdamFrames=200M, sticky actions=true, backbone=ResNet v2, Layer Normalisation=true, optimizer=Adam2021.04 | 1,006.4 | 2,856.2 | — | — | — | — | |
| MuZeroFrames=200M2021.04 | 741.7 | 2,183.6 | — | — | — | — | |
| MuZero stickyFrames=200M, sticky actions=true2021.04 | 692.9 | 2,188.4 | — | — | — | — | |
| LASERFrames=200M2021.04 | 431 | — | — | — | — | — | |
| UNREALFrames=250M, Hyper-parameters=tuned per game2021.04 | 250 | 880 | — | — | — | — | |
| RainbowFrames=200M2021.04 | 231.1 | — | — | — | — | — | |
| QR-DQN-1Training frames=200 million, Evaluation protocol=best agent protocol, Random no-ops=up to 30, Exploration rate (epsilon)=0.0012017.10 | 211 | 915 | 41 | 54 | — | — | |
| QR-DQN-0Training frames=200 million, Evaluation protocol=best agent protocol, Random no-ops=up to 30, Exploration rate (epsilon)=0.0012017.10 | 199 | 881 | 38 | 52 | — | — | |
| IMPALAFrames=200M2021.04 | 191.8 | 957.6 | — | — | — | — | |
| C51Training frames=200 million, Evaluation protocol=best agent protocol, Random no-ops=up to 30, Exploration rate (epsilon)=0.0012017.10 | 178 | 701 | 40 | 50 | — | — | |
| PR. DUEL.Training frames=200 million, Evaluation protocol=best agent protocol, Random no-ops=up to 30, Exploration rate (epsilon)=0.0012017.10 | 172 | 592 | 39 | 44 | — | — | |
| DUEL.Training frames=200 million, Evaluation protocol=best agent protocol, Random no-ops=up to 30, Exploration rate (epsilon)=0.0012017.10 | 151 | 373 | 37 | 50 | — | — | |
| PRIOR.Training frames=200 million, Evaluation protocol=best agent protocol, Random no-ops=up to 30, Exploration rate (epsilon)=0.0012017.10 | 124 | 434 | 39 | 48 | — | — | |
| DDQNTraining frames=200 million, Evaluation protocol=best agent protocol, Random no-ops=up to 30, Exploration rate (epsilon)=0.0012017.10 | 118 | 307 | 33 | 43 | — | — | |
| DQNTraining frames=200 million, Evaluation protocol=best agent protocol, Random no-ops=up to 30, Exploration rate (epsilon)=0.0012017.10 | 79 | 228 | 24 | 0 | — | — | |
| A3CTraining Time=4 days, Resources (per game)=16 cores2018.03 | — | — | — | — | — | 117 | |
| Ape-X DQNTraining Time=5 days, Environment Frames=22800M, Resources (per game)=376 cores, 1 GPU (Tesla P100)2018.03 | — | — | — | — | 434 | 358 | |
| Distributional (C51)Training Time=10 days, Environment Frames=200M, Resources (per game)=1 GPU2018.03 | — | — | — | — | 178 | 125 | |
| DQNTraining Time=9.5 days, Environment Frames=200M, Resources (per game)=1 GPU2018.03 | — | — | — | — | 79 | 68 | |
| Gorila DQNTraining Time=~4 days, Resources (per game)=>100 CPUs, with a mixed number of cores per CPU machine2018.03 | — | — | — | — | 96 | 78 | |
| Prioritized DuelingTraining Time=9.5 days, Environment Frames=200M, Resources (per game)=1 GPU2018.03 | — | — | — | — | 172 | 115 | |
| RainbowTraining Time=10 days, Environment Frames=200M, Resources (per game)=1 GPU2018.03 | — | — | — | — | 223 | 153 | |
| UNREALEnvironment Frames=250M, Resources (per game)=16 cores, Evaluation Subset=49 games, Hyperparameter Tuning=Tuned per game2018.03 | — | — | — | — | 331 | 250 |