Optimization Modeling on NL4OPT (Accuracy and Runtime)
96.3AccuracySIRL-Qwen2.5-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SIRL-Qwen2.5-7BMethod Category=Learning-based Methods2026.06 | 96.3 | — | |
| OptMATH-Qwen2.5-7BMethod Category=Learning-based Methods2026.06 | 94.7 | — | |
| StarORMethod Category=Test-Time Methods, Base model=Qwen3-4B-Instruct-25072026.06 | 92.7 | — | |
| GPT-4Method Category=Model-based Methods2026.06 | 89 | — | |
| OR-R1-Qwen3-8BMethod Category=Test-Time Methods, Base model=Qwen3-4B-Instruct-25072026.06 | 88.3 | — | |
| SAC-Opt2025.09 | 86.8 | 78.43 | |
| SAC-OptBackbone=GPT-4o, Number of runs=52025.09 | 86.8 | — | |
| SAC-Opt2025.09 | 86.8 | — | |
| OptiTreeMethod Category=Test-Time Methods, Base model=Qwen3-4B-Instruct-25072026.06 | 86.4 | — | |
| ORLM-LLaMA3-8BMethod Category=Learning-based Methods2026.06 | 85.7 | — | |
| DeepSeek-V3.1Method Category=Model-based Methods2026.06 | 84.8 | — | |
| DeepSeek-R1Method Category=Model-based Methods2026.06 | 82.4 | — | |
| LLMOPT-Qwen2.5-14BMethod Category=Learning-based Methods2026.06 | 80.3 | — | |
| Best-baseline2025.09 | 79.8 | 64.67 | |
| OptiMUS-0.3Backbone=GPT-4o, Number of runs=52025.09 | 79.8 | — | |
| OptiMUSversion=0.32025.09 | 79.8 | — | |
| AutoFormulatorMethod Category=Test-Time Methods, Base model=Qwen3-4B-Instruct-25072026.06 | 75.5 | — | |
| Best-of-NMethod Category=Test-Time Methods, Base model=Qwen3-4B-Instruct-2507, N=162026.06 | 70.6 | — | |
| OpenAI-o3Method Category=Model-based Methods2026.06 | 69.4 | — | |
| OptiMUS-0.2Backbone=GPT-4o, Number of runs=52025.09 | 69.2 | — | |
| OptiMUSversion=0.22025.09 | 69.2 | — | |
| ReflexionBackbone=GPT-4o, Number of runs=52025.09 | 68.2 | — | |
| Reflexion2025.09 | 68.2 | — | |
| CAFAreferenced=true2025.09 | 68.1 | — | |
| CoEreferenced=true2025.09 | 66.7 | — | |
| CoTreferenced=true2025.09 | 62.2 | — | |
| Standardreferenced=true2025.09 | 61.2 | — | |
| ReflexionMethod Category=Test-Time Methods, Base model=Qwen3-4B-Instruct-25072026.06 | 58 | — | |
| Zero-shotMethod Category=Test-Time Methods, Base model=Qwen3-4B-Instruct-25072026.06 | 55.9 | — | |
| CAFA2025.09 | — | 7.52 | |
| CoE2025.09 | — | 69.68 | |
| CoT2025.09 | — | 7.55 | |
| OptiMUSthreshold=0.22025.09 | — | 59.41 | |
| OptiMUSthreshold=0.32025.09 | — | 64.67 | |
| Reflexion2025.09 | — | 8.32 | |
| SAC-Opt-LLMverification=LLM-based2025.09 | — | 78.43 | |
| SAC-Opt-Simverification=similarity-based2025.09 | — | 198.82 | |
| Standard2025.09 | — | 5.3 |