Policy Gradient on MuJoCo Hopper (final 20 iterations)
185.9Average ReturnHybrid (λ*)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Hybrid (λ*)lambda=optimal2025.11 | 185.9 | 12.6 | -0.81 | |
| Pathwise2025.11 | 177.7 | 10.8 | -0.54 | |
| REINFORCE2025.11 | 116.3 | 7 | — |
| Method | Links | |||
|---|---|---|---|---|
| Hybrid (λ*)lambda=optimal2025.11 | 185.9 | 12.6 | -0.81 | |
| Pathwise2025.11 | 177.7 | 10.8 | -0.54 | |
| REINFORCE2025.11 | 116.3 | 7 | — |