Function Calling on BFCL (Accuracy)
77.9AccuracyLlama 3.1 Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama 3.1 InstructModel Scale=70B2025.04 | 77.9 | |
| ParamΔModel Scale=70B2025.04 | 77.7 | |
| Llama 3 InstructModel Scale=70B2025.04 | 76.8 | |
| Qwen3-14BGroup=Larger, Configuration=Optimal serving and decoding2026.03 | 74.3 | |
| Qwen3-30B-A3BGroup=Larger, Configuration=Optimal serving and decoding2026.03 | 74.1 | |
| Llama 3.1 InstructModel Scale=8B2025.04 | 67.9 | |
| ParamΔModel Scale=8B2025.04 | 60.9 | |
| Llama 3 InstructModel Scale=8B2025.04 | 60.1 | |
| Gpt-oss-20b-highGroup=Larger, Configuration=Optimal serving and decoding2026.03 | 58.9 | |
| gemma-3-12b-itGroup=Larger, Configuration=Optimal serving and decoding2026.03 | 52.2 | |
| EngGPT2-16B-A3BGroup=Comparable, Configuration=Optimal serving and decoding2026.03 | 48.5 | |
| gemma-2-9b-itGroup=Comparable, Configuration=Optimal serving and decoding2026.03 | 44.4 | |
| Moonlight-16B-A3B-InstructGroup=Comparable, Configuration=Optimal serving and decoding2026.03 | 42.2 | |
| Llama-3.1-8B-InstructGroup=Comparable, Configuration=Optimal serving and decoding2026.03 | 37.4 |