Continuous Control on Ant (Value Estimation MAE and Normalized Return)
15.8Value Estimation MAEStandard RL w/ DR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Standard RL w/ DRRL Algorithm=Standard A2C, OPE Baseline=DR2025.09 | 15.8 | — | |
| EvA-RL w/ EvaluatorRL Algorithm=EvA-RL, Evaluator=co-learned, Beta (β)=5 × 10−42025.09 | 16.47 | 99.9 | |
| Standard RL w/ PDISRL Algorithm=Standard A2C, OPE Baseline=PDIS2025.09 | 28.38 | — | |
| Standard RL w/ FQERL Algorithm=Standard A2C, OPE Baseline=FQE2025.09 | 46.58 | — |