Reinforcement Learning on Atari Pong
21Mean Episode ReturnR2D2
Evaluation Results
| Method | Links | |
|---|---|---|
| R2D22020.02 | 21 | |
| R2D2Loss=Retrace2020.02 | 20.9 | |
| R2D2Extension=RND2020.02 | 20.7 | |
| NGUN=322020.02 | 19.6 | |
| OPRIDEObservation=pixel-based, Human-in-the-loop=true2026.02 | 17.8 | |
| IDRLObservation=pixel-based, Human-in-the-loop=true2026.02 | 15.3 | |
| Human2020.02 | 14.6 | |
| OPRLObservation=pixel-based, Human-in-the-loop=true2026.02 | 9.6 | |
| PTObservation=pixel-based, Human-in-the-loop=true2026.02 | 9.4 | |
| PT+PDSObservation=pixel-based, Human-in-the-loop=true2026.02 | 8.5 | |
| Human2025.04 | -3 | |
| MTSpark_ADD# Trainable Parameters=3,300,3572025.04 | -5.4 | |
| NGUN=1, Variant=without RND2020.02 | -8.1 | |
| SwitchMT# Trainable Parameters=3,300,3572025.04 | -8.8 | |
| NGUN=12020.02 | -9.4 | |
| DSQN# Trainable Parameters=1,693,6822025.04 | -11.2 | |
| DSQN_D# Trainable Parameters=3,300,3392025.04 | -13.8 | |
| DQN# Trainable Parameters=1,693,6822025.04 | -18.6 | |
| DQN_D# Trainable Parameters=3,300,3392025.04 | -20.2 |