Agent Planning and API Calling on ToolBench Out of the domain
83.9Plan AccuracyGPT-4o
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| GPT-4o2025.06 | 83.9 | 63.2 | 42.9 | 22.3 | 44.1 | 99.8 | 59.3 | |
| DeepSeek-V32025.06 | 77.2 | 56.1 | 32 | 21.9 | 38.2 | 100 | 54.2 | |
| DeepSeek-R12025.06 | 74.6 | 52 | 28.3 | 20.4 | 35.6 | 99.8 | 51.8 | |
| HiMA-R1Model Parameters=3B+3B, Thinking Process=Enabled2025.06 | 73.4 | 57.5 | 31.1 | 37 | 48.9 | 96.9 | 57.5 | |
| HiMA-SFTModel Parameters=7B+3B, Thinking Process=Enabled2025.06 | 72.3 | 57 | 32.3 | 36.1 | 48.7 | 96.3 | 57.1 | |
| HiMA-SFT-noModel Parameters=7B+3B, Thinking Process=Disabled2025.06 | 59.7 | 41.1 | 19.7 | 25.9 | 36.5 | 97.3 | 46.7 | |
| Qwen2.5-14BModel Parameters=14B2025.06 | 43.9 | 16.6 | 7.2 | 6.2 | 12.3 | 99.9 | 31 | |
| Qwen2.5-32BModel Parameters=32B2025.06 | 28.6 | 0 | 0 | 0 | 0 | 100 | 21.4 |