Control on LunarLander (Robustness Gap)
0.18Robustness GapUP(mlp(wide))
Evaluation Results
| Method | Links | |
|---|---|---|
| UP(mlp(wide))Context=GRAVITY_Y, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 0.18 | |
| UP(gru)Context=GRAVITY_X, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 0.21 | |
| UP(lstm)Context=GRAVITY_X, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 0.24 | |
| UP(mlp(narrow))Context=GRAVITY_Y, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 0.32 | |
| UP(lstm)Context=GRAVITY_Y, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 0.37 | |
| UP(gru)Context=GRAVITY_Y, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 0.55 | |
| UP(lstm)Context=G_X,G_Y, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 0.56 | |
| UP(mlp(narrow))Context=GRAVITY_X, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 0.93 | |
| UP(mlp(narrow))Context=G_X,G_Y, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 1.09 | |
| UP(gru)Context=G_X,G_Y, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 1.28 | |
| UP(mlp(wide))Context=G_X,G_Y, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 1.4 | |
| UP(mlp(wide))Context=GRAVITY_X, Training Algorithm=PPO, State-action history length (k)=4, Data retention fraction (alpha)=Tuned2026.04 | 1.47 |