Goal-conditioned Reinforcement Learning on Modified PandaReachDense End-effector space, +N-G v3
-2.25Avg RewardP2S
Evaluation Results
| Method | Links | |
|---|---|---|
| P2SAction Space=End-effector, Initial Joint Randomization=true, Goal Randomization=false2025.04 | -2.25 | |
| O2SAction Space=End-effector, Initial Joint Randomization=true, Goal Randomization=false2025.04 | -2.36 | |
| Q2SAction Space=End-effector, Initial Joint Randomization=true, Goal Randomization=false2025.04 | -2.47 | |
| RSAction Space=End-effector, Initial Joint Randomization=true, Goal Randomization=false2025.04 | -3.62 |