Reinforcement Learning on D4RL Cheetah Random
91.9Mean Normalized ReturnBest
Evaluation Results
| Method | Links | |
|---|---|---|
| BestPhase=Offline, Online Interaction Budget=160 K2026.05 | 91.9 | |
| OEPhase=Offline, Online Interaction Budget=160 K2026.05 | 91.8 | |
| Adaptive Policy Selection and Fine-TuningPhase=Online, Online Interaction Budget=160 K2026.05 | 91.7 | |
| FTPhase=Online, Online Interaction Budget=160 K2026.05 | 76.5 | |
| OPEPhase=Offline, Online Interaction Budget=160 K2026.05 | 55.8 |