Optimization Modeling and Solving on NL4Opt
95.22Solution AccuracyMiniOpt-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| MiniOpt-7BCategory=Ours, Avg.=64.76, Rank=1, Rank*=12026.06 | 95.22 | |
| MiniOpt-3BCategory=Ours, Avg.=59.65, Rank=5, Rank*=22026.06 | 93.04 | |
| DeepSeek-R1 (671B)Category=General Models (Thinking), Avg.=60.85, Rank=22026.06 | 83.91 | |
| GPT-5Category=General Models (Thinking), Avg.=57.54, Rank=62026.06 | 80.43 | |
| LLMOPT-14BCategory=Learning-based Models, Avg.=60.10, Rank=42026.06 | 80.28 | |
| OptMATH-7BCategory=Learning-based Models, Avg.=54.62, Rank=8, Rank*=32026.06 | 78.7 | |
| DeepSeek-V3 (671B)Category=General Models, Avg.=60.14, Rank=32026.06 | 78.26 | |
| Gemini-2.5-ProCategory=General Models (Thinking), Avg.=57.39, Rank=72026.06 | 78.26 | |
| Step-OPT-Qwen2.5-7BCategory=Learning-based Models, Avg.=52.22, Rank=9, Rank*=42026.06 | 77.83 | |
| Qwen2.5-14B-InstructCategory=General Models, Avg.=47.46, Rank=102026.06 | 67.39 | |
| Chain-of-ExpertsCategory=Prompt-based Methods, Avg.=45.78, Rank=112026.06 | 66.52 | |
| ReflexionCategory=Prompt-based Methods, Avg.=45.54, Rank=122026.06 | 56.52 | |
| Qwen2.5-7B-InstructCategory=General Models, Avg.=33.20, Rank=14, Rank*=62026.06 | 53.48 | |
| Step-OPT-Qwen2.5-3BCategory=Learning-based Models, Avg.=39.76, Rank=13, Rank*=52026.06 | 41.3 | |
| Qwen3-8BCategory=General Models (Thinking), Avg.=21.79, Rank=16, Rank*=72026.06 | 30.87 | |
| Qwen3-14BCategory=General Models (Thinking), Avg.=23.78, Rank=152026.06 | 24.35 | |
| Qwen2.5-3B-InstructCategory=General Models, Avg.=11.23, Rank=18, Rank*=92026.06 | 19.13 | |
| Qwen3-4BCategory=General Models (Thinking), Avg.=11.16, Rank=19, Rank*=82026.06 | 16.52 | |
| OptiMUSCategory=Prompt-based Methods, Avg.=20.65, Rank=172026.06 | 13.48 |