Reinforcement Learning on HalfCheetah-vel online downstream setting
2.8Normalized RewardTask-Only
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Task-OnlyEvaluation Protocol=Online Downstream2026.05 | 2.8 | 1 | |
| OracleEvaluation Protocol=Online Downstream2026.05 | 1 | 0 | |
| Safe-VPLEvaluation Protocol=Online Downstream2026.05 | 0.96 | 0.004 | |
| SOPLEvaluation Protocol=Online Downstream2026.05 | 0.9 | 0.007 | |
| Safe-CPLEvaluation Protocol=Online Downstream2026.05 | 0.89 | 0.003 | |
| RCEvaluation Protocol=Online Downstream, omega=0.52026.05 | 0.82 | 0.076 |