Offline Reinforcement Learning on D4RL hopper-medium-expert-sparse
113Average Normalized RewardGORMPO
Evaluation Results
| Method | Links | |
|---|---|---|
| GORMPOBase Model=MOBILE, Density Estimator=VAE2026.05 | 113 | |
| GORMPOBase Model=MOBILE, Density Estimator=DDPM2026.05 | 113 | |
| GORMPOBase Model=MOBILE, Density Estimator=RealNVP2026.05 | 95.5 | |
| MOBILE2026.05 | 95.2 | |
| MOBILEBase Model=MOBILE2026.05 | 95.2 | |
| GORMPOBase Model=MOBILE, Density Estimator=KDE2026.05 | 95.2 | |
| SPOT2026.05 | 83.8 | |
| SPOT2026.05 | 83.8 | |
| GORMPODensity Estimator=DDPM2026.05 | 14.1 | |
| GORMPOBase Model=MBPO, Density Estimator=DDPM2026.05 | 14.1 | |
| GORMPODensity Estimator=RealNVP2026.05 | 10 | |
| GORMPOBase Model=MBPO, Density Estimator=RealNVP2026.05 | 10 | |
| GORMPODensity Estimator=KDE2026.05 | 8.91 | |
| GORMPOBase Model=MBPO, Density Estimator=KDE2026.05 | 8.91 | |
| GORMPODensity Estimator=NeuralODE2026.05 | 5.05 | |
| MBPO2026.05 | 4.81 | |
| MBPOBase Model=MBPO2026.05 | 4.81 | |
| GORMPODensity Estimator=VAE2026.05 | 2.32 | |
| GORMPOBase Model=MBPO, Density Estimator=VAE2026.05 | 2.32 |