Agentic Performance on TAU2-Bench
85.4Success RateGemini 3-Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 3-Pro2026.02 | 85.4 | |
| DS V3.2-Thinking2026.02 | 80.3 | |
| GPT-5 (High)2026.02 | 80.1 | |
| ERNIE 5.02026.02 | 78.79 | |
| Lowest CentroidModel=Minimax-M2.5, Model Scale=230BA10B2026.04 | 73 | |
| Pass@1Model=Minimax-M2.5, Model Scale=230BA10B2026.04 | 66.1 | |
| Greedy DecodingModel=Minimax-M2.5, Model Scale=230BA10B2026.04 | 65 | |
| Lowest CentroidModel=Qwen3-Coder, Model Scale=480BA35B2026.04 | 65 | |
| Pass@1Model=Qwen3-Coder, Model Scale=480BA35B2026.04 | 58.5 | |
| Gemini 2.5-Pro2026.02 | 56.2 | |
| Greedy DecodingModel=Qwen3-Coder, Model Scale=480BA35B2026.04 | 56 | |
| Bottom WindowModel=Qwen3-Coder, Model Scale=480BA35B2026.04 | 53 | |
| Tail ConfidenceModel=Qwen3-Coder, Model Scale=480BA35B2026.04 | 50 | |
| Bottom WindowModel=Minimax-M2.5, Model Scale=230BA10B2026.04 | 48 | |
| Tail ConfidenceModel=Minimax-M2.5, Model Scale=230BA10B2026.04 | 32 | |
| Greedy DecodingModel=Qwen3-Next, Model Scale=80BA3B2026.04 | 27 | |
| Lowest CentroidModel=Qwen3-Next, Model Scale=80BA3B2026.04 | 27 | |
| Pass@1Model=Qwen3-Next, Model Scale=80BA3B2026.04 | 20.1 | |
| Bottom WindowModel=Qwen3-Next, Model Scale=80BA3B2026.04 | 9 | |
| Tail ConfidenceModel=Qwen3-Next, Model Scale=80BA3B2026.04 | 2 |