Reinforcement Learning on D4RL Walker (Medium-Replay)
96.3Mean Normalized ReturnAdaptive Policy Selection and Fine-Tuning
Evaluation Results
| Method | Links | |
|---|---|---|
| Adaptive Policy Selection and Fine-TuningPhase=Online, Online Interaction Budget=160 K2026.05 | 96.3 | |
| BestPhase=Offline, Online Interaction Budget=160 K2026.05 | 95.8 | |
| OEPhase=Offline, Online Interaction Budget=160 K2026.05 | 95.3 | |
| OPEPhase=Offline, Online Interaction Budget=160 K2026.05 | 34.1 | |
| FTPhase=Online, Online Interaction Budget=160 K2026.05 | 28.7 |