Optimization Modeling and Solving on NLP4LP
89.02SA ScoreORThought
Evaluation Results
| Method | Links | |
|---|---|---|
| ORThought2026.07 | 89.02 | |
| OptiAgentBackbone=GPT 5.42026.07 | 87.12 | |
| LLMOPT2026.07 | 83.8 | |
| GPT 5.2System Prompt=with2026.07 | 83.46 | |
| Sonnet 4.5System Prompt=with2026.07 | 81.82 | |
| GPT 5.4System Prompt=with2026.07 | 81.61 | |
| NEMO2026.07 | 81.4 | |
| GPT 5.4System Prompt=without2026.07 | 80.46 | |
| DeepSeek-V3 (671B)Category=General Models, Avg.=60.14, Rank=32026.06 | 79.34 | |
| SonnetSystem Prompt=without2026.07 | 78.41 | |
| DeepSeek-R1System Prompt=with2026.07 | 76.89 | |
| OptiAgentBackbone=Sonnet 4.52026.07 | 76.14 | |
| CoE2026.07 | 75 | |
| MiniOpt-7BCategory=Ours, Avg.=64.76, Rank=1, Rank*=12026.06 | 74.79 | |
| MiniOpt-3BCategory=Ours, Avg.=59.65, Rank=5, Rank*=22026.06 | 74.38 | |
| GPT 5.2System Prompt=without2026.07 | 74.24 | |
| Gemini-2.5-ProCategory=General Models (Thinking), Avg.=57.39, Rank=72026.06 | 73.55 | |
| LLMOPT-14BCategory=Learning-based Models, Avg.=60.10, Rank=42026.06 | 73.42 | |
| GPT-5Category=General Models (Thinking), Avg.=57.54, Rank=62026.06 | 73.14 | |
| DeepSeek-R1 (671B)Category=General Models (Thinking), Avg.=60.85, Rank=22026.06 | 69.83 | |
| OptMATH-7BCategory=Learning-based Models, Avg.=54.62, Rank=8, Rank*=32026.06 | 68.6 | |
| ORLM2026.07 | 67.8 | |
| Qwen2.5-14B-InstructCategory=General Models, Avg.=47.46, Rank=102026.06 | 67.36 | |
| DeepSeekSystem Prompt=without2026.07 | 64.02 | |
| Chain-of-ExpertsCategory=Prompt-based Methods, Avg.=45.78, Rank=112026.06 | 59.09 | |
| Qwen2.5-7B-InstructCategory=General Models, Avg.=33.20, Rank=14, Rank*=62026.06 | 55.37 | |
| ReflexionCategory=Prompt-based Methods, Avg.=45.54, Rank=122026.06 | 53.72 | |
| Step-OPT-Qwen2.5-3BCategory=Learning-based Models, Avg.=39.76, Rank=13, Rank*=52026.06 | 53.31 | |
| Step-OPT-Qwen2.5-7BCategory=Learning-based Models, Avg.=52.22, Rank=9, Rank*=42026.06 | 48.35 | |
| Qwen3-8BCategory=General Models (Thinking), Avg.=21.79, Rank=16, Rank*=72026.06 | 34.71 | |
| Qwen3-14BCategory=General Models (Thinking), Avg.=23.78, Rank=152026.06 | 19.42 | |
| Qwen2.5-3B-InstructCategory=General Models, Avg.=11.23, Rank=18, Rank*=92026.06 | 18.6 | |
| OptiMUSCategory=Prompt-based Methods, Avg.=20.65, Rank=172026.06 | 18.18 | |
| Qwen3-4BCategory=General Models (Thinking), Avg.=11.16, Rank=19, Rank*=82026.06 | 15.29 |