Offline Reinforcement Learning on D4RL walker2d-medium-expert-sparse
118Average Normalized RewardGORMPO
Evaluation Results
| Method | Links | |
|---|---|---|
| GORMPOBase Model=MOBILE, Density Estimator=RealNVP2026.05 | 118 | |
| GORMPOBase Model=MOBILE, Density Estimator=KDE2026.05 | 117 | |
| GORMPOBase Model=MOBILE, Density Estimator=VAE2026.05 | 117 | |
| GORMPOBase Model=MOBILE, Density Estimator=DDPM2026.05 | 117 | |
| MOBILE2026.05 | 116 | |
| MOBILEBase Model=MOBILE2026.05 | 116 | |
| SPOT2026.05 | 113.5 | |
| SPOT2026.05 | 113.5 | |
| GORMPODensity Estimator=RealNVP2026.05 | 10.1 | |
| GORMPOBase Model=MBPO, Density Estimator=RealNVP2026.05 | 10.1 | |
| GORMPODensity Estimator=VAE2026.05 | 8.2 | |
| GORMPOBase Model=MBPO, Density Estimator=VAE2026.05 | 8.2 | |
| GORMPODensity Estimator=KDE2026.05 | 6.31 | |
| GORMPOBase Model=MBPO, Density Estimator=KDE2026.05 | 6.31 | |
| GORMPODensity Estimator=DDPM2026.05 | 5.23 | |
| GORMPOBase Model=MBPO, Density Estimator=DDPM2026.05 | 5.23 | |
| GORMPODensity Estimator=NeuralODE2026.05 | 3.44 | |
| MBPO2026.05 | 3.19 | |
| MBPOBase Model=MBPO2026.05 | 3.19 |