Reinforcement Learning on Atari 2600 Qbert
684,700ScoreNGU
Evaluation Results
| Method | Links | |
|---|---|---|
| NGUN=12020.02 | 684,700 | |
| NGUN=1, Variant=without RND2020.02 | 647,100 | |
| NGUN=322020.02 | 465,800 | |
| R2D2Loss=Retrace2020.02 | 415,600 | |
| R2D22020.02 | 408,800 | |
| R2D2Extension=RND2020.02 | 353,500 | |
| BYOL-ExplorePure-exploration regime=true, Episodes=10, Seeds=32022.06 | 200,081.13 | |
| OPRIDEObservation=pixel-based, Human-in-the-loop=true2026.02 | 13,535.2 | |
| Average HumanPure-exploration regime=true2022.06 | 13,455 | |
| Human2020.02 | 13,400 | |
| RNDPure-exploration regime=true, Episodes=10, Seeds=32022.06 | 12,233.18 | |
| IDRLObservation=pixel-based, Human-in-the-loop=true2026.02 | 8,294.1 | |
| OPRLObservation=pixel-based, Human-in-the-loop=true2026.02 | 7,924.6 | |
| PTObservation=pixel-based, Human-in-the-loop=true2026.02 | 7,482.9 | |
| PT+PDSObservation=pixel-based, Human-in-the-loop=true2026.02 | 6,844.2 | |
| NSRA-ESinput type=Atari RAM, # of neurons=~6502018.06 | 1,350 | |
| IDVQ+DRSC+XNESinput type=raw pixel, max run length=200 interactions, frameskip=5, # of neurons=62018.06 | 1,250 | |
| ICMPure-exploration regime=true, Episodes=10, Seeds=32022.06 | 1,062.49 | |
| HyperNeatinput type=raw pixel, # of neurons=~30342018.06 | 695 | |
| OpenAI ESinput type=raw pixel, # of neurons=~6502018.06 | 147.5 |