Failure attribution on τ-bench
75.9Agent AccuracyOur Baseline
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Our BaselineBackbone LLM=GPT-52026.02 | 75.9 | 32.2 | |
| Who&When*Backbone LLM=GPT-5, Prompt-modified=true, Variant=Best-performing2026.02 | 62 | 17.2 |
| Method | Links | ||
|---|---|---|---|
| Our BaselineBackbone LLM=GPT-52026.02 | 75.9 | 32.2 | |
| Who&When*Backbone LLM=GPT-5, Prompt-modified=true, Variant=Best-performing2026.02 | 62 | 17.2 |