Policy Gradient on MuJoCo Walker2d (final 20 iterations)
214Average ReturnPathwise
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Pathwise2025.11 | 214 | 63.4 | -15 | |
| Hybrid (λ*)lambda=optimal2025.11 | 150.8 | 73.1 | -33 | |
| REINFORCE2025.11 | 137.7 | 55.1 | — |
| Method | Links | |||
|---|---|---|---|---|
| Pathwise2025.11 | 214 | 63.4 | -15 | |
| Hybrid (λ*)lambda=optimal2025.11 | 150.8 | 73.1 | -33 | |
| REINFORCE2025.11 | 137.7 | 55.1 | — |