Agentic Tool Use on τ²-Bench Retail
90.4AccuracySeed2.0 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Seed2.0 Pro2026.06 | 90.4 | |
| Claude-Opus-4.52026.06 | 88.9 | |
| Claude-Sonnet-4.52026.06 | 86.2 | |
| Gemini-3-pro High2026.06 | 85.3 | |
| Qwen3.5-27BArchitecture=Dense, # Total Params=27B, # Activated Params=27B, Reasoning Mode=REASONING2026.04 | 84.7 | |
| GPT-5.2 High2026.06 | 82 | |
| K-EXAONE-236B-A23BArchitecture=MoE, # Total Params=236B, # Activated Params=23B, Reasoning Mode=REASONING2026.04 | 78.6 | |
| GPT-5 miniReasoning Mode=REASONING: HIGH2026.04 | 78.3 | |
| EXAONE 4.5 33BArchitecture=Dense, # Total Params=33B, # Activated Params=32B, Reasoning Mode=REASONING2026.04 | 77.9 | |
| Qwen3-VL-235B-A22BArchitecture=MoE, # Total Params=236B, # Activated Params=22B, Reasoning Mode=Thinking2026.04 | 67 |