Tool-use Agent Task on τ2-bench (full)
82.46Retail Success RateGPT-5.1 thinking
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GPT-5.1 thinkingMode=thinking2026.06 | 82.46 | 72 | 79.27 | 1,520,619 | 17.52 | |
| WRIT-4BBackbone=Qwen3-4B-Instruct-25072026.06 | 71.05 | 61 | 67.99 | 251,405 | — | |
| GPT-5.1 no-thinkMode=no-think2026.06 | 69.3 | 48 | 62.8 | 318,180 | 5.56 |