Agent Execution on EnterpriseBench (test)
55Execution AccuracyClaude-3.5-Sonnet
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude-3.5-SonnetPrompting=2-shot2026.03 | 55 | |
| Gemini-2.5 ProPrompting=2-shot2026.03 | 55 | |
| Qwen3-8B Agentic GRPOTraining Strategy=Agentic GRPO, Training Data=<1K2026.03 | 51 | |
| GPT-4oPrompting=2-shot2026.03 | 47 | |
| ToolAceTraining Data=26K-trained2026.03 | 41 | |
| xLAM-2-70BTraining Data=60K-trained2026.03 | 40 | |
| Qwen3-8B SFTTraining Strategy=SFT, Training Data=<1K2026.03 | 38 | |
| Qwen3-8B BasePrompting=2-shot2026.03 | 35 |