Optimization modeling and solving on IndustryOR (Pass@1 (SA))
69.1Pass@1 (SA)DeepSeek-V3
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-V3Params=671B, Category=Zero-shot LLMs2026.06 | 69.1 | |
| GPT-4oParams=Closed, Category=Zero-shot LLMs2026.06 | 66.7 | |
| Qwen2.5-72B-InstParams=72B, Category=Zero-shot LLMs2026.06 | 61.9 | |
| Qwen3-32BParams=32B, Category=Zero-shot LLMs2026.06 | 57.1 | |
| EvoOptiGraphParams=8B, Category=Fine-tuned LLMs2026.06 | 57.1 | |
| OptiMUS-v0.3Params=Closed, Category=Agentic Methods2026.06 | 54.3 | |
| ORLMParams=8B, Category=Fine-tuned LLMs2026.06 | 42.9 | |
| CoTParams=Closed, Category=Agentic Methods2026.06 | 40.5 | |
| OptSkills-DCategory=Skill-Based Methods, Base Model=DeepSeek-V3.22026.05 | 36 | |
| Gemini-3.1-ProCategory=General Models2026.05 | 34 | |
| Trace2SkillCategory=Skill-Based Methods2026.05 | 34 | |
| ORMindCategory=Agent-Based Methods2026.05 | 32 | |
| AlphaOPTCategory=Experience-Enhanced Methods2026.05 | 32 | |
| DeepSeek-R1 (671B)Category=General Models (Thinking), Avg.=60.85, Rank=22026.06 | 32 | |
| CoEParams=Closed, Category=Agentic Methods2026.06 | 31.2 | |
| ORThoughtCategory=Agent-Based Methods2026.05 | 31 | |
| DeepSeek-V3.2Category=General Models2026.05 | 30 | |
| GPT-5.4Category=General Models2026.05 | 29 | |
| Qwen3-235BCategory=General Models2026.05 | 29 | |
| LLMOPT-14BCategory=Learning-based Models, Avg.=60.10, Rank=42026.06 | 29 | |
| LLMOPT (origin)Params=14B, Category=Fine-tuned LLMs2026.06 | 29 | |
| Chain-of-ExpertsCategory=Agent-Based Methods2026.05 | 28 | |
| LEAN-LLM-OPTCategory=Experience-Enhanced Methods2026.05 | 28 | |
| Gemini-2.5-ProCategory=General Models (Thinking), Avg.=57.39, Rank=72026.06 | 28 | |
| OptSkills-QCategory=Skill-Based Methods, Base Model=Qwen3-235B-A22b-instruct-25072026.05 | 27 | |
| Step-OPT-Qwen2.5-7BCategory=Learning-based Models, Avg.=52.22, Rank=9, Rank*=42026.06 | 27 | |
| OptiMUSCategory=Agent-Based Methods2026.05 | 26 | |
| DeepSeek-V3 (671B)Category=General Models, Avg.=60.14, Rank=32026.06 | 26 | |
| GPT-5Category=General Models (Thinking), Avg.=57.54, Rank=62026.06 | 26 | |
| MiniOpt-7BCategory=Ours, Avg.=64.76, Rank=1, Rank*=12026.06 | 25 | |
| Qwen2.5-14B-InstructCategory=General Models, Avg.=47.46, Rank=102026.06 | 22 | |
| Step-OPT-Qwen2.5-3BCategory=Learning-based Models, Avg.=39.76, Rank=13, Rank*=52026.06 | 21 | |
| Chain-of-ExpertsCategory=Prompt-based Methods, Avg.=45.78, Rank=112026.06 | 19 | |
| ReflexionCategory=Prompt-based Methods, Avg.=45.54, Rank=122026.06 | 19 | |
| OptMATH-7BCategory=Learning-based Models, Avg.=54.62, Rank=8, Rank*=32026.06 | 19 | |
| MiniOpt-3BCategory=Ours, Avg.=59.65, Rank=5, Rank*=22026.06 | 17 | |
| Qwen3-14BCategory=General Models (Thinking), Avg.=23.78, Rank=152026.06 | 15 | |
| Qwen2.5-7B-InstructCategory=General Models, Avg.=33.20, Rank=14, Rank*=62026.06 | 13 | |
| OptiMUSCategory=Prompt-based Methods, Avg.=20.65, Rank=172026.06 | 8 | |
| Qwen3-8BCategory=General Models (Thinking), Avg.=21.79, Rank=16, Rank*=72026.06 | 6 | |
| Qwen2.5-3B-InstructCategory=General Models, Avg.=11.23, Rank=18, Rank*=92026.06 | 2 | |
| Qwen3-4BCategory=General Models (Thinking), Avg.=11.16, Rank=19, Rank*=82026.06 | 2 |