Offline Reinforcement Learning on D4RL halfcheetah-medium-expert-sparse
106Average Normalized RewardGORMPO
Evaluation Results
| Method | Links | |
|---|---|---|
| GORMPOBase Model=MOBILE, Density Estimator=VAE2026.05 | 106 | |
| MOBILEBase Model=MOBILE2026.05 | 104 | |
| GORMPOBase Model=MOBILE, Density Estimator=RealNVP2026.05 | 104 | |
| GORMPOBase Model=MOBILE, Density Estimator=KDE2026.05 | 103.7 | |
| GORMPOBase Model=MOBILE, Density Estimator=DDPM2026.05 | 103 | |
| MOBILE2026.05 | 99.4 | |
| GORMPODensity Estimator=VAE2026.05 | 89.9 | |
| GORMPOBase Model=MBPO, Density Estimator=VAE2026.05 | 89.9 | |
| GORMPODensity Estimator=RealNVP2026.05 | 86 | |
| GORMPOBase Model=MBPO, Density Estimator=RealNVP2026.05 | 86 | |
| GORMPODensity Estimator=DDPM2026.05 | 85.7 | |
| GORMPOBase Model=MBPO, Density Estimator=DDPM2026.05 | 85.7 | |
| GORMPODensity Estimator=NeuralODE2026.05 | 80.3 | |
| GORMPODensity Estimator=KDE2026.05 | 78.4 | |
| GORMPOBase Model=MBPO, Density Estimator=KDE2026.05 | 78.4 | |
| SPOT2026.05 | 76.4 | |
| SPOT2026.05 | 76.4 | |
| MBPO2026.05 | 63.7 | |
| MBPOBase Model=MBPO2026.05 | 63.7 |