Function Calling on BFCL (Complexity Metrics)
83.27Success Rate (Simple)OpenFunctions-v2
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| OpenFunctions-v2Type=Function Calling, Size=7B2025.10 | 83.27 | 93 | 85.5 | 66 | 81.94 | |
| Hammer2.1-3BType=Function Calling, Size=3B2025.10 | 81.42 | 95 | 89.5 | 81.5 | 86.86 | |
| Hammer2.1-3B + ToolPRM (Ours)Type=Inference Scaling, Size=3B2025.10 | 80.5 | 95.5 | 91.5 | 88 | 88.88 | |
| Qwen2.5-72B-InstructType=General Purpose, Size=72B2025.10 | 80.25 | 97.5 | 93.5 | 92 | 90.81 | |
| GPT-4o-miniType=General Purpose, Size=-2025.10 | 80.08 | 90.5 | 89.5 | 87 | 86.77 | |
| Hammer2.1-7B + MajorityType=Inference Scaling, Size=7B2025.10 | 79.58 | 95 | 93.5 | 85 | 88.27 | |
| Hammer2.1-3B + MajorityType=Inference Scaling, Size=3B2025.10 | 79.5 | 95 | 89 | 81 | 86.13 | |
| Hammer2.1-7B + ToolPRM (Ours)Type=Inference Scaling, Size=7B2025.10 | 79.08 | 95.5 | 94.5 | 89 | 89.52 | |
| Hammer2.1-1.5B + ToolPRM (Ours)Type=Inference Scaling, Size=1.5B2025.10 | 78.42 | 92 | 87.5 | 84.5 | 85.61 | |
| Hammer2.1-7BType=Function Calling, Size=7B2025.10 | 78.08 | 95 | 93.5 | 88 | 88.65 | |
| Hammer2.1-1.5B + MajorityType=Inference Scaling, Size=1.5B2025.10 | 77.5 | 91.5 | 84.5 | 80 | 83.38 | |
| GPT-4oType=General Purpose, Size=-2025.10 | 77.17 | 95 | 93.5 | 85 | 87.67 | |
| Hammer2.1-3B + Best of N (ORM)Type=Inference Scaling, Size=3B2025.10 | 77.08 | 93.5 | 89.5 | 84 | 86.02 | |
| Mistral-Nemo-InstructType=General Purpose, Size=12B2025.10 | 77 | 93.5 | 89.5 | 84.5 | 86.13 | |
| Hammer2.1-7B + Best of N (ORM)Type=Inference Scaling, Size=7B2025.10 | 76.58 | 94.5 | 91.5 | 87 | 87.4 | |
| Qwen2.5-7B-InstructType=General Purpose, Size=7B2025.10 | 75.33 | 94.5 | 91.5 | 84.5 | 86.46 | |
| Hammer2.1-1.5B + Best of N (ORM)Type=Inference Scaling, Size=1.5B2025.10 | 75.25 | 90 | 86.5 | 84 | 83.94 | |
| Hammer2.1-1.5BType=Function Calling, Size=1.5B2025.10 | 74.67 | 92 | 84.5 | 80 | 82.79 | |
| Hammer2.1-7B + Token-level Beam SearchType=Inference Scaling, Size=7B2025.10 | 74.58 | 93.5 | 91.5 | 82.5 | 85.52 | |
| Hammer2.1-3B + Token-level Beam SearchType=Inference Scaling, Size=3B2025.10 | 74.5 | 91 | 87.5 | 77 | 82.5 | |
| Qwen2.5-3B-InstructType=General Purpose, Size=3B2025.10 | 74.17 | 90.5 | 79.5 | 79 | 80.79 | |
| xLAM-7b-fc-rType=Function Calling, Size=7B2025.10 | 73.08 | 93.5 | 87 | 84 | 84.4 | |
| Llama-3.1-8B-InstructType=General Purpose, Size=8B2025.10 | 72.83 | 93.5 | 87 | 83.5 | 84.21 | |
| Qwen2.5-32B-InstructType=General Purpose, Size=32B2025.10 | 72.83 | 94 | 93.5 | 88.5 | 87.21 | |
| GRANITEType=Function Calling, Size=20B2025.10 | 72.83 | 91.5 | 84 | 81.5 | 82.46 | |
| Qwen2.5-1.5B-InstructType=General Purpose, Size=1.5B2025.10 | 72.42 | 87 | 81.5 | 75.5 | 79.11 | |
| Hammer2.1-1.5B + Token-level Beam SearchType=Inference Scaling, Size=1.5B2025.10 | 72.33 | 91.5 | 81 | 73.5 | 79.58 | |
| xLAM-1b-fc-rType=Function Calling, Size=1.3B2025.10 | 69.67 | 89.5 | 79 | 66.5 | 76.17 | |
| Hammer2.1-0.5BType=Function Calling, Size=0.5B2025.10 | 68 | 83 | 71.5 | 54 | 69.13 |