ResearchBenchmarksReinforcement Learning on HalfCheetah HurdleFollow4,790.1Cumulative RewardDTSemNets68.77041,294.50022,520.233,745.9598May 18, 2026Evaluation ResultsMethodMethodLinksCumulative RewardDTSemNets2026.054,790.1VIPER (PPO)2026.051,107.49π-PRLPolicy Type=DiscretizedPolicy Type=Discretized2026.05848.84PPO2026.05780.43DiPRL2026.05723.88π-PRLPolicy Type=Fine-tunedPolicy Type=Fine-tuned2026.05443.8π-PRLPolicy Type=Continuous...Policy Type=Continuous relaxed2026.05250.36