Multi-turn tool calling on τ2-bench
38Airline ScoreDiGiT-TC
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| DiGiT-TCBackbone=Llama3.1-8B-Instruct2026.01 | 38 | 17.62 | 6.1 | 8.77 | — | |
| Simia2026.04 | 38 | — | 35.1 | 3.5 | 25.5 | |
| Llama3.1-8B-InstructModel Type=Base model2026.01 | 30 | 17.02 | 6.14 | 14.91 | — | |
| AWM2026.04 | 30 | — | 36.8 | 12.3 | 26.4 | |
| Tool-N12026.04 | 28 | — | 36.8 | 10.5 | 25.1 | |
| TRUSTEE2026.04 | 26 | — | 45.6 | 17.5 | 29.7 | |
| ToolACE-MT-8BBackbone=Llama3.1-8B-Instruct, Note=Results reflect authors' running of their publicly released model2026.01 | 24 | 13.56 | 7.02 | 9.65 | — | |
| QWEN3-8B2026.04 | 24 | — | 43 | 15.8 | 27.6 | |
| ToolRL2026.04 | 22 | — | 16.7 | 14 | 17.6 | |
| EnvScaler2026.04 | 22 | — | 4.4 | 9.7 | 12 | |
| ToucanBackbone=Qwen2.5-7B-Instruct2026.01 | 20 | 17.77 | 22.8 | 10.5 | — | |
| Qwen2.5-7B-InstructModel Type=Base model2026.01 | 14 | 16.08 | 17.54 | 16.7 | — | |
| Best-of-N Selection (4o-s5-4o-v2)Base Model=GPT-4o, Mechanism=Best-of-N Selection (sN), Reviewer Model=GPT-4o, N=5, Prompt Version=v22026.04 | 0.48 | — | 0.614 | 0.482 | 0.525 | |
| Best-of-N Grading (4o-g5-4o-v1)Base Model=GPT-4o, Mechanism=Best-of-N Grading (gN), Reviewer Model=GPT-4o, N=5, Prompt Version=v12026.04 | 0.473 | — | 0.605 | 0.453 | 0.51 | |
| Best-of-N Grading (4o-g5-4o-v2)Base Model=GPT-4o, Mechanism=Best-of-N Grading (gN), Reviewer Model=GPT-4o, N=5, Prompt Version=v22026.04 | 0.467 | — | 0.658 | 0.488 | 0.538 | |
| Best-of-N Selection (4o-s5-4o-v1)Base Model=GPT-4o, Mechanism=Best-of-N Selection (sN), Reviewer Model=GPT-4o, N=5, Prompt Version=v12026.04 | 0.427 | — | 0.611 | 0.38 | 0.473 | |
| 4o baselineBase Model=GPT-4o, Mechanism=Baseline2026.04 | 0.42 | — | 0.629 | 0.412 | 0.487 | |
| Progressive Feedback (4o-r5-4o-v1)Base Model=GPT-4o, Mechanism=Progressive Feedback (rN), Reviewer Model=GPT-4o, N=5, Prompt Version=v12026.04 | 0.407 | — | 0.626 | 0.64 | 0.558 | |
| Progressive Feedback (4o-r5-4o-v2)Base Model=GPT-4o, Mechanism=Progressive Feedback (rN), Reviewer Model=GPT-4o, N=5, Prompt Version=v22026.04 | 0.407 | — | 0.585 | 0.596 | 0.529 |