Reinforcement Learning on HalfCheetah discrete-regime, exponential dwell
18,598Mean ReturnBAPR (full)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| BAPR (full)Number of seeds (N)=5, training iterations=1 500, c_penalty=0.5, K (test modes)=4, evaluation window=last 200 iterations, per-transition belief storage=Enabled2026.05 | 18,598 | 7,872 | — | |
| BAPR (no per-trans)Number of seeds (N)=5, training iterations=1 500, c_penalty=0.5, K (test modes)=4, evaluation window=last 200 iterations, per-transition belief storage=Disabled2026.05 | 17,862 | 3,327 | 14 | |
| ESCPNumber of seeds (N)=5, training iterations=1 500, c_penalty=0.5, K (test modes)=4, evaluation window=last 200 iterations2026.05 | 17,297 | 3,356 | — | |
| SACNumber of seeds (N)=5, training iterations=1 500, c_penalty=0.5, K (test modes)=4, evaluation window=last 200 iterations2026.05 | 15,682 | 4,785 | — |