Multi-step tool-calling on ACEBench
0.85End-to-End Success RateGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5Generation type=N/A2026.06 | 0.85 | 94.3 | |
| Qwen3-4B-Think + Full (1000)Generation type=connected generation with tool results (✓)2026.06 | 0.8 | 92 | |
| Qwen3-4B-ThinkingGeneration type=N/A2026.06 | 0.7 | 88.8 | |
| Qwen3-4B-Think + LoRA (3689)Generation type=offline generation without tool feedback (×)2026.06 | 0.65 | 83.6 | |
| Qwen3-4B-Think + Full (3689)Generation type=offline generation without tool feedback (×)2026.06 | 0.65 | 76.3 | |
| Qwen3-4B-InstructGeneration type=N/A2026.06 | 0.15 | 25.1 |