Reinforcement Learning on D4RL Cheetah Medium-Expert
98.7Mean Normalized ReturnBest
Evaluation Results
| Method | Links | |
|---|---|---|
| BestPhase=Offline, Online Interaction Budget=160 K2026.05 | 98.7 | |
| OEPhase=Offline, Online Interaction Budget=160 K2026.05 | 97.8 | |
| Adaptive Policy Selection and Fine-TuningPhase=Online, Online Interaction Budget=160 K2026.05 | 97.4 | |
| FTPhase=Online, Online Interaction Budget=160 K2026.05 | 51 | |
| OPEPhase=Offline, Online Interaction Budget=160 K2026.05 | 40.3 |