ResearchBenchmarksGoal-conditioned Reinforcement Learning on SMAXFollow95.6IQMCPPO83.53686.66889.892.932May 13, 2026Evaluation ResultsMethodMethodLinksIQMCPPONumber of tasks=6 task...Number of tasks=6 tasks, Action space=discrete, Agent type=multi-agent2026.0595.6ICSACNumber of tasks=6 task...Number of tasks=6 tasks, Action space=discrete, Agent type=multi-agent2026.0584