Reinforcement Learning on Ant-vel online downstream setting
1.11Normalized RewardTask-Only
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Task-OnlyEvaluation Protocol=Online Downstream2026.05 | 1.11 | 1 | |
| OracleEvaluation Protocol=Online Downstream2026.05 | 1 | 0 | |
| Safe-VPLEvaluation Protocol=Online Downstream2026.05 | 0.94 | 0.001 | |
| RCEvaluation Protocol=Online Downstream, omega=0.52026.05 | 0.9 | 0.1 | |
| Safe-CPLEvaluation Protocol=Online Downstream2026.05 | 0.88 | 0.028 | |
| SOPLEvaluation Protocol=Online Downstream2026.05 | 0.84 | 0.039 |