Reinforcement Learning on D4RL Ant Medium-Expert
99.1Mean Normalized ReturnBest
Evaluation Results
| Method | Links | |
|---|---|---|
| BestPhase=Offline, Online Interaction Budget=160 K2026.05 | 99.1 | |
| OEPhase=Offline, Online Interaction Budget=160 K2026.05 | 98.7 | |
| Adaptive Policy Selection and Fine-TuningPhase=Online, Online Interaction Budget=160 K2026.05 | 96.6 | |
| OPEPhase=Offline, Online Interaction Budget=160 K2026.05 | 47.8 | |
| FTPhase=Online, Online Interaction Budget=160 K2026.05 | 19.4 |