Tool Use (Pass@1) on Tau-Bench
85.4Pass@1Gemini-3.0 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-3.0 Proformat=standard function call, thinking mode=true2025.12 | 85.4 | |
| Claude-4.5-Sonnetformat=standard function call, thinking mode=true2025.12 | 84.7 | |
| DeepSeek-V3.2format=standard function call, thinking mode=true2025.12 | 80.3 | |
| GPT-5 Highformat=standard function call, thinking mode=true2025.12 | 80.2 | |
| MiniMax M2format=standard function call, thinking mode=true2025.12 | 76.9 | |
| Kimi-K2format=standard function call, thinking mode=true2025.12 | 74.3 |