ResearchBenchmarksMulti-agent reinforcement learning on Multi-soft-robot coordinationFollow108,000Convergence Episodes (x10^5)TD-MAPPO99,600156,300213,000269,700May 1, 2026Evaluation ResultsMethodMethodLinksConvergence Episodes (x10^5)Sample EfficiencyRobustness ScoreTD-MAPPO2026.05108,0002.2892TD-MASAC2026.05157,0001.9589HAPPO2026.05162,0001.2476MAAC2026.05176,0001.5272FACMAC2026.05181,0001.2176TD-MADDPG2026.05200,0001.7386MASAC2026.05217,0001.2874MADDPG2026.05262,0000.9567MAPPO2026.05270,0001.0871QDDPG2026.05318,0000.8463