Offline Reinforcement Learning on D4RL Hopper random (Mean Normalized Score)
62.1Mean Normalized ScoreAdaptive Policy Selection and Fine-Tuning
Evaluation Results
| Method | Links | |
|---|---|---|
| Adaptive Policy Selection and Fine-TuningPhase=Online, Online Interaction Budget=160 K2026.05 | 62.1 | |
| FTPhase=Online, Online Interaction Budget=160 K2026.05 | 57.6 | |
| VIPOMethod Category=Conservative model-based2025.12 | 33.4 | |
| ADMPOMethod Category=Conservative model-based2025.12 | 32.7 | |
| LEQMethod Category=Conservative model-based2025.12 | 32.4 | |
| MOBILEMethod Category=Conservative model-based2025.12 | 31.9 | |
| MOPOMethod Category=Conservative model-based2025.12 | 31.7 | |
| CBOPMethod Category=Bayesian-inspired2025.12 | 31.4 | |
| APE-VMethod Category=Conservative model-based2025.12 | 31.3 | |
| SUMOMethod Category=Conservative model-based2025.12 | 30.8 | |
| EDACMethod Category=Model-free2025.12 | 25.3 | |
| NEUBAYMethod Category=Ours2025.12 | 24.5 | |
| RAMBOMethod Category=Model-free2025.12 | 21.6 | |
| MoMoMethod Category=Conservative model-based2025.12 | 18.3 | |
| COMBOMethod Category=Model-free2025.12 | 17.9 | |
| MAPLEMethod Category=Bayesian-inspired2025.12 | 10.7 | |
| MoDAPMethod Category=Bayesian-inspired2025.12 | 8.9 | |
| CQLMethod Category=Model-free2025.12 | 5.3 | |
| BestPhase=Offline, Online Interaction Budget=160 K2026.05 | 3.8 | |
| OEPhase=Offline, Online Interaction Budget=160 K2026.05 | 3.8 | |
| OPEPhase=Offline, Online Interaction Budget=160 K2026.05 | 0.6 |