Function Calling on Tool-Alpaca
77.66F1 ScoreGPT-4o
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4oType=General Purpose, Size=-2025.10 | 77.66 | 88.64 | 66.67 | |
| Mistral-Nemo-InstructType=General Purpose, Size=12B2025.10 | 76.75 | 87.31 | 66.18 | |
| Qwen2.5-72B-InstructType=General Purpose, Size=72B2025.10 | 74.39 | 83.21 | 65.56 | |
| Qwen2.5-32B-InstructType=General Purpose, Size=32B2025.10 | 73.47 | 85.82 | 61.11 | |
| Hammer2.1-7B + ToolPRM (Ours)Type=Inference Scaling, Size=7B2025.10 | 73.36 | 81.42 | 65.3 | |
| Hammer2.1-7B + Best of N (ORM)Type=Inference Scaling, Size=7B2025.10 | 73.3 | 78.86 | 67.73 | |
| Hammer2.1-1.5B + ToolPRM (Ours)Type=Inference Scaling, Size=1.5B2025.10 | 72.93 | 82.68 | 63.18 | |
| Hammer2.1-7BType=Function Calling, Size=7B2025.10 | 72.77 | 80.93 | 64.6 | |
| Hammer2.1-7B + MajorityType=Inference Scaling, Size=7B2025.10 | 72.15 | 79.03 | 65.26 | |
| Hammer2.1-3B + ToolPRM (Ours)Type=Inference Scaling, Size=3B2025.10 | 71.96 | 80.78 | 63.13 | |
| Qwen2.5-7B-InstructType=General Purpose, Size=7B2025.10 | 71.85 | 83.7 | 60 | |
| Hammer2.1-3BType=Function Calling, Size=3B2025.10 | 71.57 | 80.31 | 62.83 | |
| Hammer2.1-7B + Token-level Beam SearchType=Inference Scaling, Size=7B2025.10 | 71.03 | 79.69 | 62.37 | |
| Hammer2.1-1.5B + Best of N (ORM)Type=Inference Scaling, Size=1.5B2025.10 | 69.54 | 77.42 | 61.66 | |
| Hammer2.1-1.5BType=Function Calling, Size=1.5B2025.10 | 69.3 | 77.42 | 61.17 | |
| Hammer2.1-0.5BType=Function Calling, Size=0.5B2025.10 | 68.89 | 77.1 | 60.67 | |
| Hammer2.1-3B + Token-level Beam SearchType=Inference Scaling, Size=3B2025.10 | 68.81 | 78.29 | 59.32 | |
| Hammer2.1-3B + Best of N (ORM)Type=Inference Scaling, Size=3B2025.10 | 68.73 | 77.29 | 60.16 | |
| Hammer2.1-1.5B + MajorityType=Inference Scaling, Size=1.5B2025.10 | 68.25 | 72.88 | 63.61 | |
| two-step training model (AugFC)Model Size=32B2026.04 | 67.75 | — | — | |
| GRANITEType=Function Calling, Size=20B2025.10 | 67.64 | 77.27 | 58 | |
| Hammer2.1-3B + MajorityType=Inference Scaling, Size=3B2025.10 | 67.14 | 76 | 58.27 | |
| two-step training model (AugFC)Model Size=7B2026.04 | 66.67 | — | — | |
| xLAM-2-32B-fcModel Size=32B2026.04 | 65.97 | — | — | |
| Llama-3.1-8B-InstructType=General Purpose, Size=8B2025.10 | 65.38 | 75.64 | 55.12 | |
| Hammer2.1-1.5B + Token-level Beam SearchType=Inference Scaling, Size=1.5B2025.10 | 64.86 | 73.03 | 56.68 | |
| Qwen2.5-32BModel Size=32B2026.04 | 64.48 | — | — | |
| xLAM-7b-fc-rType=Function Calling, Size=7B2025.10 | 63.11 | 67.26 | 58.96 | |
| Hammer-7BModel Size=7B2026.04 | 62.5 | — | — | |
| two-step training model (AugFC)Model Size=1.5B2026.04 | 62.33 | — | — | |
| OpenFunctions-v2Type=Function Calling, Size=7B2025.10 | 62.1 | 72.93 | 51.26 | |
| Qwen2.5-7BModel Size=7B2026.04 | 61.58 | — | — | |
| Qwen2.5-3B-InstructType=General Purpose, Size=3B2025.10 | 61.22 | 70.8 | 51.63 | |
| GPT-4o-miniType=General Purpose, Size=-2025.10 | 59.52 | 64.34 | 54.69 | |
| xLAM-7B-fcModel Size=7B2026.04 | 58.96 | — | — | |
| xLAM-1b-fc-rType=Function Calling, Size=1.3B2025.10 | 57.72 | 64.86 | 50.58 | |
| Hammer-1.5BModel Size=1.5B2026.04 | 53.48 | — | — | |
| Qwen2.5-1.5B-InstructType=General Purpose, Size=1.5B2025.10 | 52.65 | 62.07 | 43.23 | |
| xLAM-1.3B-fcModel Size=1.5B2026.04 | 50.58 | — | — | |
| Qwen2.5-1.5BModel Size=1.5B2026.04 | 42.9 | — | — |