Multi-Agent Reinforcement Learning on CAGE Challenge 4 (test)
8,144R̄ACD3-GAT
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| ACD3-GATGroup=Safety, Ep.=600, S=32026.06 | 8,144 | 11,074 | 13.8 | 48.2 | 0.3 | |
| Rule-basedGroup=Nonlearn., Ep.=140, S=32026.06 | 8,124 | 10,456 | 100 | 115.9 | 0 | |
| ACD3+AskHumanGroup=ACD3 comp., Ep.=30, S=12026.06 | 7,901 | 10,460 | 10 | 36.8 | 0 | |
| ACD3+det. G-CRPGroup=ACD3 comp., Ep.=30, S=12026.06 | 7,901 | 10,460 | 10 | 36.8 | 0 | |
| C-MAPPO-GATGroup=Safety, Ep.=600, S=32026.06 | 6,992 | 9,694 | 0.3 | 15.5 | 0 | |
| SleepGroup=Nonlearn., Ep.=140, S=32026.06 | 6,792 | 8,984 | 0 | 0 | 0 | |
| Fact.-IPPOGroup=Arch., Ep.=100, S=12026.06 | 5,378 | 7,428 | 100 | 311.1 | 12 | |
| RandomGroup=Nonlearn., Ep.=140, S=32026.06 | 5,149 | 6,737 | 100 | 426.1 | 10.7 | |
| CVaR-MAPPOGroup=Safety, Ep.=100, S=12026.06 | 5,131 | 6,509 | 100 | 429.6 | 9 | |
| IA2CGroup=Actor-critic, Ep.=100, S=12026.06 | 4,948 | 6,691 | 100 | 420.6 | 7 | |
| Opp-IPPOGroup=Arch., Ep.=100, S=12026.06 | 4,648 | 7,015 | 100 | 321.2 | 0 | |
| MAPPO-GNNGroup=MAPPO enc., Ep.=100, S=12026.06 | 4,387 | 6,108 | 100 | 406.2 | 1 | |
| IPPOGroup=Actor-critic, Ep.=600, S=32026.06 | 4,254 | 6,299 | 100 | 314.4 | 1.2 | |
| MAPPO-GATGroup=MAPPO enc., Ep.=600, S=32026.06 | 3,979 | 5,864 | 100 | 355.4 | 3.5 | |
| MAPPO-MLPGroup=MAPPO enc., Ep.=200, S=12026.06 | 3,937 | 6,226 | 100 | 316.2 | 1.5 |