Optimization Modeling on LogiOR OptiBench (out-of-distribution)
56LogiOR ScoreAlphaOPT
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| AlphaOPTCategory=Experience-learning, Backbone=gpt-4o2025.10 | 56 | 93.5 | 86.5 | 74.8 | |
| AlphaOPTCategory=Ablation, Refinement=false2025.10 | 48.9 | 92.1 | 84.1 | 70.5 | |
| ExpelCategory=Experience-learning, Backbone=gpt-4o2025.10 | 47.8 | 80.4 | 74.3 | 64.1 | |
| StandardCategory=Prompt-based, Backbone=gpt-4o2025.10 | 46.7 | 72.7 | 67.9 | 59.7 | |
| AlphaOPTCategory=Ablation, Insight Example=false2025.10 | 45.7 | 91 | 82.9 | 68.4 | |
| ORThoughtCategory=Prompt-based, Backbone=gpt-4o2025.10 | 44.6 | 84.4 | 77 | 64.5 | |
| ReflexionCategory=Experience-learning, Backbone=gpt-4o2025.10 | 43.5 | 76.9 | 70.7 | 60.2 | |
| LLMOPTCategory=Fine-tuning, Backbone=Qwen2.5-14B2025.10 | 40.2 | 66.4 | 61.5 | 53.3 | |
| ORLMCategory=Fine-tuning, Backbone=LLaMa3-8B2025.10 | 19.6 | 78.2 | 67.3 | 48.9 | |
| OptiMusCategory=Prompt-based, Backbone=gpt-4o2025.10 | 17.4 | 74.7 | 64.1 | 46 |