Multi-agent Strategy Generation on Repeated Rock-Paper-Scissors final iteration (K = 20)
201Population ReturnLLM Agent (70B)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LLM Agent (70B)Agent Parameters=70B2026.03 | 201 | 45.8 | 155.2 | |
| LLM Agent (27B)Agent Parameters=27B2026.03 | 193.2 | 67.2 | 126 | |
| ContRM2026.03 | 164.8 | 16.3 | 148.5 | |
| CSRO - LinearRef. (code)Oracle=LinearRefinement, Input Format=code, K=20, M=10, Backbone=Gemini 2.5 Pro2026.03 | 159.8 | 37.7 | 122.1 | |
| CSRO - ZeroShot (desc.)Oracle=ZeroShot, Input Format=description, K=20, Backbone=Gemini 2.5 Pro2026.03 | 130.2 | 66.7 | 63.5 | |
| PSRO-IMPALA2026.03 | 108.9 | 423.2 | 532.1 | |
| CSRO - LinearRef. (desc.)Oracle=LinearRefinement, Input Format=description, K=20, M=10, Backbone=Gemini 2.5 Pro2026.03 | 99.3 | 31.6 | 67.7 | |
| CSRO - AlphaEvolveOracle=AlphaEvolve, K=20, Backbone=Gemini 2.5 Pro2026.03 | 50.5 | 25.2 | 25.4 | |
| QLRecall Length=10 rounds2026.03 | 0.5 | 8.6 | 9.1 |