Tool Use Reasoning on τ-Bench
63.9Avg Accuracyo1
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| o1Model Category=Proprietary Model2026.02 | 63.9 | 73.5 | 54.2 | |
| Claude-3.7-SonnetModel Category=Proprietary Model2026.02 | 59.8 | 78.3 | 41.2 | |
| GPT-4oModel Category=Proprietary Model2026.02 | 52.9 | 62.8 | 43 | |
| D-CORE-14BModel Category=Custom-Trained Model, Backbone=Qwen3-14B, Parameters=14B2026.02 | 51.3 | 56.5 | 46 | |
| xLAM2-70BModel Category=Open-Source Model, Parameters=70B2026.02 | 48.3 | 57.7 | 38.8 | |
| DeepSeek-R1Model Category=Proprietary Model2026.02 | 47.8 | 55.6 | 40 | |
| D-CORE-8BModel Category=Custom-Trained Model, Backbone=Qwen3-8B, Parameters=8B2026.02 | 47.6 | 50.7 | 44.4 | |
| xLAM2-32BModel Category=Open-Source Model, Parameters=32B2026.02 | 45 | 52.5 | 37.6 | |
| xLAM2-8BModel Category=Open-Source Model, Parameters=8B2026.02 | 42.4 | 50.7 | 34 | |
| ToolRL-Qwen3-14BModel Category=Custom-Trained Model, Backbone=Qwen3-14B2026.02 | 37.8 | 49.5 | 26 | |
| Qwen3-32BModel Category=Open-Source Model, Parameters=32B2026.02 | 36.6 | 39.6 | 33.6 | |
| Qwen3-14BModel Category=Open-Source Model, Parameters=14B2026.02 | 33.6 | 41.6 | 25.6 | |
| ToolRL-Qwen3-8BModel Category=Custom-Trained Model, Backbone=Qwen3-8B2026.02 | 31.5 | 40.6 | 22.4 | |
| Qwen3-8BModel Category=Open-Source Model, Parameters=8B2026.02 | 29 | 34.7 | 23.2 |