Mathematical formulation of design requirements on D_HQ (test)
80.12AobjLLAMA3.1-8B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LLAMA3.1-8BModel Category=Ours (APF) based on open-source LLMs, Fine-tuned=APF2025.12 | 80.12 | 79.69 | 79.76 | |
| Qwen2.5-7BModel Category=Ours (APF) based on open-source LLMs, Fine-tuned=APF2025.12 | 79.9 | 79.59 | 79.61 | |
| Mistral-7BModel Category=Ours (APF) based on open-source LLMs, Fine-tuned=APF2025.12 | 79.74 | 78.83 | 79.18 | |
| Chain-of-ExpertsModel Category=Baselines2025.12 | 74.26 | 74.53 | 72.52 | |
| DeepSeek-V3Model Category=Baselines2025.12 | 74.04 | 76.9 | 75.18 | |
| OptimusModel Category=Baselines2025.12 | 63.41 | 69.86 | 66.87 | |
| GPT-4oModel Category=Baselines2025.12 | 60.55 | 70.75 | 66.51 | |
| Qwen2.5-7BModel Category=Open-Source LLMs2025.12 | 35.42 | 73.33 | 52.92 | |
| Mistral-7BModel Category=Open-Source LLMs2025.12 | 7.33 | 49.36 | 30.07 | |
| LLAMA3.1-8BModel Category=Open-Source LLMs2025.12 | -4.53 | 50.29 | 22.48 |