Continuous Control on HalfCheetah (Value Estimation MAE and Normalized Return)
4.18Value Estimation MAEEvA-RL w/ Evaluator
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| EvA-RL w/ EvaluatorRL Algorithm=EvA-RL, Evaluator=co-learned, Beta (β)=5 × 10−42025.09 | 4.18 | 1.052 | |
| Standard RL w/ PDISRL Algorithm=Standard A2C, OPE Baseline=PDIS2025.09 | 5.28 | — | |
| Standard RL w/ DRRL Algorithm=Standard A2C, OPE Baseline=DR2025.09 | 7.29 | — | |
| Standard RL w/ FQERL Algorithm=Standard A2C, OPE Baseline=FQE2025.09 | 29.19 | — |