Tool Use on BFCL
94AccuracyOfficial Instruct Model
Evaluation Results
| Method | Links | |
|---|---|---|
| Official Instruct ModelEvaluation Category=Official Instruct Model2026.05 | 94 | |
| Opus-4.7 (xHigh)Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 91.33 | |
| GLM-4.7 & ANDESEvaluation Category=Proposed Method Comparison, Base Model=Qwen3-1.7B2026.05 | 89 | |
| Opus-4.6 (1M)Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 87.33 | |
| Gemini-3.1-ProEvaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 85.67 | |
| Qwen3-4BParameters=4B2026.05 | 68.5 | |
| Qwen3-8BParameters=8B2026.05 | 68.5 | |
| Qwen 3 VL 32B InstructParameters=32B2025.12 | 66.3 | |
| Qwen 3 VL 8B Inststage=Instruct2025.12 | 66.2 | |
| GPT-5 nanoScale=nano2026.05 | 64.4 | |
| Qwen 3 32BThinking=No, Parameters=32B2025.12 | 63.1 | |
| Qwen 2.5 32BParameters=32B2025.12 | 62.8 | |
| Ministral-3-8BParameters=8B2026.05 | 62 | |
| Qwen 3 8Bstage=Instruct2025.12 | 60.2 | |
| Olmo 3.1 32B InstructStage=Final Instruct 3.12025.12 | 58.8 | |
| Olmo 3.1 32B InstructStage=DPO2025.12 | 58.6 | |
| gemma-3-12b-itParameters=12B2026.05 | 58.2 | |
| Olmo 3.1 32B InstructStage=SFT2025.12 | 57 | |
| Qwen 2.5 7Bstage=Instruct2025.12 | 55.8 | |
| gpt-oss-20bParameters=20B2026.05 | 54.6 | |
| Llama-3.1-8BParameters=8B2026.05 | 54.2 | |
| Olmo 3 7B Instructstage=Final Instruct2025.12 | 49.8 | |
| Olmo 3 7B Instructstage=DPO2025.12 | 49.6 | |
| EngGPT2-16B-A3BParameters=16B2026.05 | 49.1 | |
| Olmo 3 7B Instructstage=SFT2025.12 | 48.9 | |
| Llama-3.2-3BParameters=3B2026.05 | 45.3 | |
| gemma-3-4b-itParameters=4B2026.05 | 43.2 | |
| GLM-4.7 (Scaffold-only)Evaluation Category=Proposed Method Comparison, Base Model=Qwen3-1.7B2026.05 | 43 | |
| LLaMAntino-3-8BParameters=8B2026.05 | 38.3 | |
| GPT-5.2Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 29.33 | |
| Opus-4.6Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 28.33 | |
| Moonlight-16B-A3BParameters=16B2026.05 | 26.3 | |
| Velvet-14BParameters=14B2026.05 | 25.7 | |
| Minerva-7BParameters=7B2026.05 | 25 | |
| FastwebMIIA-7BParameters=7B2026.05 | 24.5 | |
| deepseek-moe-16bParameters=16B2026.05 | 22.5 | |
| Base Model (Qwen3-1.7B)Evaluation Category=Zero-Shot2026.05 | 0 | |
| Sonnet-4.5Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0 | |
| MiniMax-M2.5Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0 | |
| MiniMax-M2.1Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0 | |
| Qwen3-MaxEvaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0 | |
| Kimi-K2-ThinkingEvaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0 | |
| GPT-5.1-Codex-MaxEvaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0 | |
| GPT-5.4 (High)Evaluation Category=Agent-Post-Trained, Base Model=Qwen3-1.7B2026.05 | 0 | |
| GLM-4.7 (OpenCode)Evaluation Category=Proposed Method Comparison, Base Model=Qwen3-1.7B2026.05 | 0 |