Agent Performance on ACEBench-en
56End-to-End AccuracyGPT-4o-2024-11-20
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4o-2024-11-202025.08 | 56 | — | 77.8 | |
| Llama3.1-70B-Inst2025.08 | 41 | — | 62.5 | |
| ToolACE-MT2025.08 | 8.4 | — | 34 | |
| Llama3.1-8B-Inst2025.08 | 6.7 | — | 18.3 | |
| Multi-Agent Simulation2025.08 | 6.7 | — | 15 | |
| ToolACE-MTAblation=Without Offline Verification2025.08 | 1.7 | — | 28.5 | |
| ToolACE-MTAblation=Without Iterative Refinement2025.08 | 1.7 | — | 22.8 | |
| DS V3.2-Thinking2026.02 | — | 81.4 | — | |
| ERNIE 5.02026.02 | — | 87.7 | — | |
| Gemini 2.5-Pro2026.02 | — | 80.9 | — | |
| Gemini 3-Pro2026.02 | — | 80.9 | — | |
| GPT-5 (High)2026.02 | — | 79.3 | — |