Reinforcement Learning on Acrobot
-82.5Average ReturnsDTSemNet
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DTSemNetNf (Number of features)=6, Na (Number of actions)=3, Height=42026.05 | -82.5 | — | |
| DGTNf (Number of features)=6, Na (Number of actions)=3, Height=42026.05 | -83.1 | — | |
| VIPERNf (Number of features)=6, Na (Number of actions)=3, Height=42026.05 | -83.92 | — | |
| Deep RLNf (Number of features)=6, Na (Number of actions)=3, Height=42026.05 | -84 | — | |
| AC-SGDbatch size=1000, seeds=52026.01 | -86.2 | — | |
| AC-Adambatch size=1000, seeds=52026.01 | -88.4 | — | |
| ICCTNf (Number of features)=6, Na (Number of actions)=3, Height=42026.05 | -88.6 | — | |
| AC-CGbatch size=1000, seeds=52026.01 | -93.3 | — | |
| SMACbatch size=1000, seeds=52026.01 | -94.1 | — | |
| AC-KFACbatch size=1000, seeds=52026.01 | -170.7 | — | |
| A2CNumber of training seeds=5, Number of evaluation episodes=1002026.03 | — | 213.6 | |
| CG-FPDNumber of training seeds=5, Number of evaluation episodes=1002026.03 | — | 90.6 | |
| DF-CWP-CPNumber of training seeds=5, Number of evaluation episodes=1002026.03 | — | 78.67 | |
| DQNNumber of training seeds=5, Number of evaluation episodes=1002026.03 | — | 61.9 | |
| PPONumber of training seeds=5, Number of evaluation episodes=1002026.03 | — | 63.5 |