Reinforcement Learning on Atari 2600 Montezuma's Revenge
18,003,200ScoreGo-Explore
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Go-ExploreAlgorithm Category=Reinforcement Learning, Domain knowledge utilization=true, Score Aggregation=Best2019.01 | 18,003,200 | — | |
| Human World RecordAlgorithm Category=Human2019.01 | 1,219,200 | — | |
| Go-ExploreAlgorithm Category=Reinforcement Learning, Domain knowledge utilization=true, Score Aggregation=Average2019.01 | 666,474 | — | |
| LfSDAlgorithm Category=Imitation Learning, Score Aggregation=Best2019.01 | 74,500 | — | |
| Go-ExploreAlgorithm Category=Reinforcement Learning, Domain knowledge utilization=false, Score Aggregation=Average2019.01 | 43,763 | — | |
| TDC+CMCAlgorithm Category=Imitation Learning2019.01 | 41,098 | — | |
| Human ExpertAlgorithm Category=Human2019.01 | 34,900 | — | |
| Ape-X DQfDAlgorithm Category=Imitation Learning2019.01 | 29,384 | — | |
| PPO+CoExAlgorithm Category=Reinforcement Learning2019.01 | 11,540 | — | |
| RNDAlgorithm Category=Reinforcement Learning2019.01 | 11,347 | — | |
| BYOL-ExplorePure-exploration regime=true, Episodes=10, Seeds=32022.06 | 5,146.73 | — | |
| Average HumanPure-exploration regime=true2022.06 | 4,753.3 | — | |
| Average HumanAlgorithm Category=Human2019.01 | 4,753 | — | |
| DQfDAlgorithm Category=Imitation Learning2019.01 | 4,638 | — | |
| DQN-CTSAlgorithm Category=Reinforcement Learning2019.01 | 3,705 | — | |
| DeepCSAlgorithm Category=Reinforcement Learning2019.01 | 3,500 | — | |
| UBEAlgorithm Category=Reinforcement Learning2019.01 | 3,000 | — | |
| DQN MMC CTSTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=52018.07 | 2,941.9 | — | |
| Feature-EBAlgorithm Category=Reinforcement Learning2019.01 | 2,745 | — | |
| ReactorAlgorithm Category=Reinforcement Learning2019.01 | 2,643.5 | — | |
| IMPALAAlgorithm Category=Reinforcement Learning2019.01 | 2,643.5 | — | |
| DQN-PixelCNNAlgorithm Category=Reinforcement Learning2019.01 | 2,514 | — | |
| Ape-XAlgorithm Category=Reinforcement Learning2019.01 | 2,500 | — | |
| RNDPure-exploration regime=true, Episodes=10, Seeds=32022.06 | 2,305.65 | — | |
| DQN MMC PIXELCNNTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=52018.07 | 1,671.7 | — | |
| DQN MMC + SRTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=52018.07 | 1,395.4 | 1,121.8 | |
| A3C-CTSAlgorithm Category=Reinforcement Learning2019.01 | 1,127 | — | |
| RNDTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=52018.07 | 524.8 | 314 | |
| SARSAAlgorithm Category=Reinforcement Learning2019.01 | 259 | — | |
| BASS-hashAlgorithm Category=Reinforcement Learning2019.01 | 238 | — | |
| RainbowAlgorithm Category=Reinforcement Learning2019.01 | 154 | — | |
| GorilaAlgorithm Category=Reinforcement Learning2019.01 | 84 | — | |
| A3CAlgorithm Category=Reinforcement Learning2019.01 | 67 | — | |
| DDQNAlgorithm Category=Reinforcement Learning2019.01 | 42 | — | |
| Duel. DQNAlgorithm Category=Reinforcement Learning2019.01 | 22 | — | |
| Prior. DQNAlgorithm Category=Reinforcement Learning2019.01 | 13 | — | |
| LinearAlgorithm Category=Reinforcement Learning2019.01 | 10.7 | — | |
| ICMPure-exploration regime=true, Episodes=10, Seeds=32022.06 | 2.68 | — | |
| DQNTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=52018.07 | 0 | 0 | |
| DQN MMCTraining frames=100 million, Sticky actions (s)=0.25, Frame skip=52018.07 | 0 | 0 | |
| DQNAlgorithm Category=Reinforcement Learning2019.01 | 0 | — | |
| MP-EBAlgorithm Category=Reinforcement Learning2019.01 | 0 | — | |
| Pop-ArtAlgorithm Category=Reinforcement Learning2019.01 | 0 | — | |
| ESAlgorithm Category=Reinforcement Learning2019.01 | 0 | — | |
| C51Algorithm Category=Reinforcement Learning2019.01 | 0 | — |