Agent Generalization on NetSecGame unseen topology (test)
95Win RateReAct
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ReActbackbone=gpt-oss-120b2026.03 | 95 | 63.86 | 31.16 | 27.76 | |
| Conceptual Q-learningTotal training episodes=331,0002026.03 | 65.53 | 62 | 67.1 | 49.8 | |
| LLM-BERT2026.03 | 51.6 | 6.5 | 57.66 | 29.16 | |
| MAMLTotal training episodes=100,0002026.03 | 40 | 50.8 | 84.8 | 61.99 | |
| Random agent (baseline)2026.03 | 6 | 100.93 | 98.47 | — | |
| Dual Buffer DQNTotal training episodes=50002026.03 | 3.07 | 104.45 | 97.82 | 15.33 | |
| Reptile (baseline)Total training episodes=100,0002026.03 | 2.76 | 105.9 | 98.94 | 58.42 | |
| Single Buffer DQNTotal training episodes=50002026.03 | 2.07 | 106.71 | 98.98 | 10.33 | |
| DDQN+emb2026.03 | 0 | -99 | 100 | — |