Operations Research Autoformalization on NL4OPT
97.3Accuracy (%)Standard
Evaluation Results
| Method | Links | |
|---|---|---|
| StandardLLM=o3, Method Category=direct prompting2024.07 | 97.3 | |
| OptMATHSize=32B2026.04 | 95.9 | |
| LLMOPTLLM=Qwen1.5-14B, Method Category=fine-tuning LLMs2024.07 | 93 | |
| OR-R1RLSize=8B2026.04 | 88.3 | |
| OptiMUS-0.3LLM=GPT-4o, Method Category=agentic frameworks2024.07 | 86.6 | |
| OR-R1SFTSize=8B2026.04 | 86 | |
| ORLMLLM=Deepseek-Math, Method Category=fine-tuning LLMs2024.07 | 85.7 | |
| Step-OptSize=8B2026.04 | 84.5 | |
| OptiMUS-0.2LLM=GPT-4o, Method Category=agentic frameworks2024.07 | 78.8 | |
| AutoORSize=8B2026.04 | 78.6 | |
| ORLMSize=8B2026.04 | 73.8 | |
| ReflexionLLM=GPT-4o, Method Category=direct prompting2024.07 | 53 | |
| StandardLLM=GPT-4o, Method Category=direct prompting2024.07 | 47.3 |