Reinforcement Learning on BipedalWalker
325.35Average Episode RewardDTSemNet
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DTSemNetNf (Number of features)=24, Na (Number of actions)=4 dim., Height=7, Variant=Top-k2026.05 | 325.35 | — | |
| Deep RLNf (Number of features)=24, Na (Number of actions)=4 dim., Height=72026.05 | 315.3 | — | |
| DTSemNetNf (Number of features)=24, Na (Number of actions)=4 dim., Height=7, Variant=STE2026.05 | 314.98 | — | |
| TD32023.11 | 314.24 | — | |
| TRPO2023.11 | 312.14 | — | |
| DSP2023.11 | 311.78 | — | |
| TD32025.02 | 310.2 | — | |
| ACKTR2023.11 | 309.57 | — | |
| ESPL2023.11 | 309.43 | — | |
| SAC2023.11 | 308.31 | — | |
| SAC2025.02 | 307.3 | — | |
| ICCTNf (Number of features)=24, Na (Number of actions)=4 dim., Height=62026.05 | 301.34 | — | |
| A2C2023.11 | 291.79 | — | |
| PPO2023.11 | 287.43 | — | |
| PPO2025.02 | 286.2 | — | |
| SALSA-RLLatent dimension size (hd)=82025.02 | 280.9 | — | |
| SALSA-RLLatent dimension size (hd)=62025.02 | 280.4 | — | |
| DSP2025.02 | 264.4 | — | |
| SALSA-RLLatent dimension size (hd)=162025.02 | 262.6 | — | |
| DGTNf (Number of features)=24, Na (Number of actions)=4 dim., Height=8, Variant=linear2026.05 | 244.5 | — | |
| A2C2025.02 | 241 | — | |
| SALSA-RLLatent dimension size (hd)=42025.02 | 235.1 | — | |
| DDPG2023.11 | 209.42 | — | |
| DDPG2025.02 | 94.2 | — | |
| DGTNf (Number of features)=24, Na (Number of actions)=4 dim., Height=8, Variant=scalar2026.05 | 78.33 | — | |
| Regression2023.11 | -110.77 | — | |
| DSP2023.11 | — | 400,000 | |
| ESPLSelection=Best policy of three independent runs2023.11 | — | 2,000 | |
| Regression2023.11 | — | 1,000 |