Combinatorial Optimization on Overall (test)
73.01Average PerformanceMEMOIR (GPT-5-mini w/ GPT-5 critic)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MEMOIR (GPT-5-mini w/ GPT-5 critic)Budget=B=162026.05 | 73.01 | 96.73 | |
| MEMOIR (GPT-5-mini)Budget=B=162026.05 | 68.84 | 92.01 | |
| FunSearchBudget=B=162026.05 | 65.7 | 87.55 | |
| ReEvoBudget=B=162026.05 | 57.93 | 79.28 | |
| GreedyRefineBudget=B=162026.05 | 56.91 | 73.43 | |
| AIDEBudget=B=162026.05 | 53.51 | 77.28 | |
| MCTS-AHDBudget=B=162026.05 | 53.32 | 80.32 | |
| o3-mini-highBudget=B=162026.05 | 44.81 | 66.22 | |
| Classical SolverBudget=B=162026.05 | 43.04 | — | |
| GPT-5-miniBudget=B=162026.05 | 38.4 | 67.84 | |
| MEMOIR (Qwen2.5-Coder-32B)Budget=B=162026.05 | 21.45 | 58.82 | |
| GPT-5 ChatBudget=B=162026.05 | 18.6 | 24.43 | |
| QwQ-32BBudget=B=162026.05 | 17.93 | 45.04 | |
| MEMOIR (Llama-3.3-70B-Instruct)Budget=B=162026.05 | 15.39 | 50.5 | |
| DeepSeek-R1-Distill-Llama-70BBudget=B=162026.05 | 13.33 | 40.57 | |
| Llama-3.3-70B-InstructBudget=B=162026.05 | 11.17 | 19.31 | |
| Qwen2.5-Coder-32BBudget=B=162026.05 | 10.73 | 13.1 |