Tool Selection on EnterpriseBench
24Tool Selection AccuracyGemini-2.5 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-2.5 ProModel Category=Closed-Source Models, Evaluation Protocol=2-shot2026.03 | 24 | |
| Claude-3.5-SonnetModel Category=Closed-Source Models, Evaluation Protocol=2-shot2026.03 | 22 | |
| GPT-4oModel Category=Closed-Source Models, Evaluation Protocol=2-shot2026.03 | 21 | |
| Qwen3-8B Agentic GRPOModel Category=Our Platform-Trained Models (<1K), Training Setting=Agentic GRPO2026.03 | 21 | |
| Qwen3-8B SFTModel Category=Our Platform-Trained Models (<1K), Training Setting=Supervised Fine-Tuning2026.03 | 17 | |
| Qwen3-8B BaseModel Category=Open-Source Models, Evaluation Protocol=2-shot2026.03 | 14 | |
| xLAM-2-70BModel Category=Open-Source Models, Training Setting=60K-trained2026.03 | 12 | |
| ToolAceModel Category=Open-Source Models, Training Setting=26K-trained2026.03 | 11 |