Agentic on τ2-Bench
91.6ScoreClaude Opus 4.5
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude Opus 4.52026.02 | 91.6 | |
| Gemini 3 Pro2026.02 | 90.7 | |
| GLM-52026.02 | 89.7 | |
| GLM-4.72026.02 | 87.4 | |
| GPT-5.2 (xhigh)2026.02 | 85.5 | |
| DeepSeek-V3.22026.02 | 85.3 | |
| Kimi K2.52026.02 | 80.2 | |
| Ling-2.6-1T2026.06 | 78.36 | |
| GLM-5Thinking Mode=non-thinking2026.06 | 78.12 | |
| Qwen3.5-Omni Flash 35B-A3BContext Length=256K2026.07 | 78 | |
| Ling-2.6-flash2026.06 | 76.36 | |
| DeepSeek-V3.2Thinking Mode=nothink2026.06 | 75.63 | |
| Kimi-K2.5Mode=Instant2026.06 | 71.21 | |
| GPT-5.4Reasoning Mode=non-reasoning2026.06 | 69.53 | |
| Nemotron-3-Super 120B-A12BEvaluation Setting=non-reasoning2026.06 | 68.92 | |
| Audex 30B-A3BContext Length=1M2026.07 | 57.2 | |
| GPT-5.4-miniEvaluation Setting=non-reasoning2026.06 | 46.92 | |
| Qwen3-Omni 30B-A3B ThinkingContext Length=64K2026.07 | 45.4 | |
| Audex 2BContext Length=128K2026.07 | 41.7 | |
| GPT-OSS-120BEvaluation Setting=low2026.06 | 23.48 |