Safe Multi-Agent Reinforcement Learning on Wireless Communication 25 agents
19.26Constraint Violation RateScalable Primal-Dual Actor-Critic
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Scalable Primal-Dual Actor-CriticAlgorithm=Ours2023.05 | 19.26 | — | |
| MAPPO-LInformation Access=Global2023.05 | 40 | — | |
| Decentralized Aggregate MAPPO-LInformation Access=Local neighborhood, Reward structure=Sum of rewards in local neighborhood2023.05 | 118.9 | — | |
| Decentralized MAPPO-LInformation Access=Local neighborhood2023.05 | 157.6 | — |