ResearchBenchmarksMulti-Objective Reinforcement Learning on B-9Follow2,500,000Environment StepsGPI-LS555,2001,060,1001,565,0002,069,900Aug 4, 2025Evaluation ResultsMethodMethodLinksEnvironment StepsHypervolume (HV)Scalarized Performance (SP)Expected Utility (EU)GPI-LS2025.082,500,000———C-MORL2025.082,500,0007.932.793.5SPFT2025.08630,0008.221.923.53