Goal-conditioned Reinforcement Learning on Modified PandaReachDense End-effector space, -N+G v3
-3.36Average RewardO2S
Evaluation Results
| Method | Links | |
|---|---|---|
| O2SAction Space=End-effector, Initial Joint Randomization=false, Goal Randomization=true2025.04 | -3.36 | |
| P2SAction Space=End-effector, Initial Joint Randomization=false, Goal Randomization=true2025.04 | -3.77 | |
| Q2SAction Space=End-effector, Initial Joint Randomization=false, Goal Randomization=true2025.04 | -3.97 | |
| RSAction Space=End-effector, Initial Joint Randomization=false, Goal Randomization=true2025.04 | -6.25 |