Multi-Agent Reinforcement Learning on smac 1o_10b_vs_1r
89Test Win RateTD
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TDCommunication Penalty (+pnlt)=false, Gating Mechanism=true, Training Step Budget Allocation=4M steps standard RL2026.07 | 89 | 93 | |
| TDCommunication Penalty (+pnlt)=false, Gating Mechanism=false, Training Step Budget Allocation=4M steps standard RL2026.07 | 87 | 97 | |
| MUTETraining Step Budget Allocation=4M total (2M pre-training, 0.5M MVE, 1.5M unlearning)2026.07 | 77 | 6 | |
| TDCommunication Penalty (+pnlt)=true, Gating Mechanism=true, Training Step Budget Allocation=4M steps standard RL2026.07 | 67 | 26 | |
| TDCommunication Penalty (+pnlt)=true, Gating Mechanism=false, Training Step Budget Allocation=4M steps standard RL2026.07 | 44 | 9 |