Reinforcement Learning on Atari Breakout
864Mean ReturnNGU
Evaluation Results
| Method | Links | |
|---|---|---|
| NGUN=1, Variant=without RND2020.02 | 864 | |
| NGUN=12020.02 | 864 | |
| R2D2Loss=Retrace2020.02 | 838.3 | |
| R2D22020.02 | 837.7 | |
| R2D2Extension=RND2020.02 | 815.8 | |
| NGUN=322020.02 | 532.8 | |
| Average Human2024.05 | 450 | |
| DDQN2024.05 | 320 | |
| DQN2024.05 | 300 | |
| FDQN2024.05 | 297 | |
| OPRIDEObservation=pixel-based, Human-in-the-loop=true2026.02 | 256.7 | |
| IDRLObservation=pixel-based, Human-in-the-loop=true2026.02 | 153.7 | |
| OPRLObservation=pixel-based, Human-in-the-loop=true2026.02 | 125.9 | |
| PT+PDSObservation=pixel-based, Human-in-the-loop=true2026.02 | 86.3 | |
| PTObservation=pixel-based, Human-in-the-loop=true2026.02 | 79.9 | |
| Human2025.04 | 31 | |
| Human2020.02 | 30.5 | |
| SwitchMT# Trainable Parameters=3,300,3572025.04 | 5.6 | |
| DQN# Trainable Parameters=1,693,6822025.04 | 3.2 | |
| MTSpark_ADD# Trainable Parameters=3,300,3572025.04 | 0.6 | |
| DSQN# Trainable Parameters=1,693,6822025.04 | 0.4 | |
| DQN_D# Trainable Parameters=3,300,3392025.04 | 0 | |
| DSQN_D# Trainable Parameters=3,300,3392025.04 | 0 |