Reinforcement Learning on D4RL Walker Medium-Expert
100Mean Normalized ReturnBest
Evaluation Results
| Method | Links | |
|---|---|---|
| BestPhase=Offline, Online Interaction Budget=160 K2026.05 | 100 | |
| OEPhase=Offline, Online Interaction Budget=160 K2026.05 | 100 | |
| Adaptive Policy Selection and Fine-TuningPhase=Online, Online Interaction Budget=160 K2026.05 | 97.5 | |
| OPEPhase=Offline, Online Interaction Budget=160 K2026.05 | 22.7 | |
| FTPhase=Online, Online Interaction Budget=160 K2026.05 | 6.1 |