Optimization Modeling on NLP4LP, NL4OPT, IndustryOR, MAMO (test)
97.3NLP4LP ScoreLLMOPT
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| LLMOPTCategory=Fine-tuning, Backbone=Qwen2.5-14B2025.10 | 97.3 | 86.5 | 44 | 85.8 | 85 | 78.4 | |
| AlphaOPTCategory=Experience-learning, Backbone=gpt-4o2025.10 | 87.7 | 81.2 | 60 | 76.5 | 80.1 | 76.4 | |
| ORLMCategory=Fine-tuning, Backbone=LLaMa3-8B2025.10 | 86.3 | 87.5 | 36 | 55.9 | 75 | 66.4 | |
| AlphaOPTCategory=Ablation, Insight Example=false2025.10 | 86.3 | 78.1 | 56 | 58.8 | 76 | 69.8 | |
| AlphaOPTCategory=Ablation, Refinement=false2025.10 | 82.2 | 80.7 | 44 | 61.8 | 73.3 | 67.2 | |
| ExpelCategory=Experience-learning, Backbone=gpt-4o2025.10 | 79.5 | 67.2 | 48 | 70.4 | 69.9 | 66.3 | |
| ReflexionCategory=Experience-learning, Backbone=gpt-4o2025.10 | 76.7 | 64.1 | 56 | 47.1 | 64.8 | 61 | |
| OptiMusCategory=Prompt-based, Backbone=gpt-4o2025.10 | 71.2 | 73.4 | 36 | 29.4 | 60.2 | 52.5 | |
| ORThoughtCategory=Prompt-based, Backbone=gpt-4o2025.10 | 69.9 | 75 | 60 | 41.2 | 65.3 | 61.5 | |
| StandardCategory=Prompt-based, Backbone=gpt-4o2025.10 | 68.5 | 54.7 | 52 | 44.1 | 57.7 | 54.8 |