ResearchBenchmarksGoal-conditioned Reinforcement Learning on ConnectorFollow95.5IQM (Normalised Win Rate)CPPO90.71691.95893.294.442May 13, 2026Evaluation ResultsMethodMethodLinksIQM (Normalised Win Rate)CPPONumber of tasks=4 task...Number of tasks=4 tasks, Action space=discrete, Agent type=multi-agent2026.0595.5ICSACNumber of tasks=4 task...Number of tasks=4 tasks, Action space=discrete, Agent type=multi-agent2026.0590.9