Task-oriented Dialogue on τ-bench 157 scenarios
45.5Collaboration SRGPT-4.1-mini
Evaluation Results
| Method | Links | ||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4.1-miniReasoning step limit=30, Number of trials=42025.09 | 45.5 | 41.7 | 39.5 | 45.1 | 45.4 | 100 | 91.6 | 86.8 | 98.9 | 99.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NCUser2025.09 | 45.5 | — | — | — | — | 97.5 | — | — | — | — | 40.9 | 98.1 | 34.6 | 96.4 | 36.8 | 99.5 | 33.8 | 97.7 | 40 | 97.7 | 38.1 | 93.8 | |
| Qwen3-235b-a22bReasoning step limit=30, Number of trials=42025.09 | 41.4 | 36.8 | 32.3 | 37.6 | 39.3 | 100 | 88.9 | 78 | 90.8 | 94.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PBUS2025.09 | 38.9 | — | — | — | — | 87.8 | — | — | — | — | 40.9 | 92.4 | 44.4 | 95.8 | 39.2 | 97.9 | 43.3 | 96 | 39.3 | 92.3 | 45.1 | 88.3 | |
| Qwen3-30b-a3bReasoning step limit=30, Number of trials=42025.09 | 27.9 | 26.6 | 20.4 | 24.8 | 30.1 | 100 | 95.3 | 73.1 | 88.9 | 107.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-3.1-70b-instructReasoning step limit=30, Number of trials=42025.09 | 21.8 | 18.5 | 14.7 | 17.8 | 16.4 | 100 | 84.9 | 67.4 | 81.7 | 75.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4.1-nanoReasoning step limit=30, Number of trials=42025.09 | 12 | 10 | 6.8 | 8.8 | 8 | 100 | 83.3 | 56.7 | 72.5 | 66.7 | — | — | — | — | — | — | — | — | — | — | — | — |