Tool-use on Tau-Bench (Pass@1 Domain Scores and Word Count)
85.5Average Pass@1ProPlay
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ProPlayBackbone LLM=GPT-4.1-mini, Temperature=0, Inference setting=Online, Trials per task=12026.06 | 85.5 | 90 | 83.5 | — | |
| Wall-EBackbone LLM=GPT-4.1-mini, Temperature=0, Inference setting=Online, Trials per task=12026.06 | 84.9 | 88 | 83.5 | — | |
| ReflexionBackbone LLM=GPT-4.1-mini, Temperature=0, Inference setting=Online, Trials per task=12026.06 | 84.2 | 90 | 81.7 | — | |
| ExpeLBackbone LLM=GPT-4.1-mini, Temperature=0, Inference setting=Online, Trials per task=12026.06 | 83 | 90 | 80 | — | |
| WorldCoderBackbone LLM=GPT-4.1-mini, Temperature=0, Inference setting=Online, Trials per task=12026.06 | 83 | 82 | 83.5 | — | |
| ReActBackbone LLM=GPT-4.1-mini, Temperature=0, Inference setting=Online, Trials per task=12026.06 | 81.8 | 84 | 80.9 | — | |
| LATSBackbone LLM=GPT-4.1-mini, Temperature=0, Inference setting=Online, Trials per task=12026.06 | 81.8 | 82 | 81.7 | — | |
| GLM-4.5Business Policy Configuration=Full Business Policy In-Context2026.03 | 67.42 | 55.35 | 79.5 | 40,300 | |
| GLM-4.5-AirBusiness Policy Configuration=Full Business Policy In-Context2026.03 | 64.91 | 54.45 | 75.37 | 42,000 | |
| Claude 3.7 SonnetBusiness Policy Configuration=Full Business Policy In-Context2026.03 | 62.12 | 44.05 | 80.2 | — | |
| Claude 3.5 SonnetBusiness Policy Configuration=Full Business Policy In-Context2026.03 | 59.35 | 48.05 | 69.54 | — | |
| Claude 4 SonnetBusiness Policy Configuration=Full Business Policy In-Context2026.03 | 59.25 | 50.7 | 68 | — | |
| PA3Business Policy Configuration=Without Business Policy, Training Stage=Stage 32026.03 | 53.75 | 42 | 65.51 | 27,000 | |
| xLAM-2-32bBusiness Policy Configuration=Full Business Policy In-Context2026.03 | 50.52 | 37.35 | 63.7 | 45,000 | |
| Claude 4.0-SonnetVersion=4.0-Sonnet2025.08 | 50.22 | — | — | — | |
| PA2Business Policy Configuration=Without Business Policy, Training Stage=Stage 22026.03 | 50.01 | 36.95 | 63.07 | 45,000 | |
| Gemini 2.5-ProVersion=2.5-Pro2025.08 | 47.09 | — | — | — | |
| Qwen-2.5-32BBusiness Policy Configuration=Full Business Policy In-Context2026.03 | 43 | 27.95 | 58.04 | 39,000 | |
| Gemini 2.5-FlashVersion=2.5-Flash2025.08 | 40.04 | — | — | — | |
| PA1Business Policy Configuration=Without Business Policy, Training Stage=Stage 12026.03 | 37.93 | 17.15 | 58.72 | 30,000 | |
| xLAM-2-32bBusiness Policy Configuration=Without Business Policy2026.03 | 37.91 | 17.4 | 58.41 | 29,700 | |
| GPT 4oVersion=4o2025.08 | 37.43 | — | — | — | |
| Qwen2.5-72B InstructFamily=Qwen2.5-72B, Version=Instruct2025.08 | 34.26 | — | — | — | |
| Qwen3-8B Reasoning FTRL-Reinforce++Family=Qwen3-8B, Version=Reasoning, Method=FTRL-Reinforce++2025.08 | 32.52 | — | — | — | |
| Qwen3-14B Reasoning FTRL-GRPOFamily=Qwen3-14B, Version=Reasoning, Method=FTRL-GRPO2025.08 | 31.7 | — | — | — | |
| Qwen3-32B ReasoningFamily=Qwen3-32B, Version=Reasoning2025.08 | 31 | — | — | — | |
| Qwen3-8B Reasoning FTRL-GRPOFamily=Qwen3-8B, Version=Reasoning, Method=FTRL-GRPO2025.08 | 28.91 | — | — | — | |
| Qwen3-32B Non-ReasoningFamily=Qwen3-32B, Version=Non-Reasoning2025.08 | 27.39 | — | — | — | |
| Qwen3-14B Reasoning FTRL-Reinforce++Family=Qwen3-14B, Version=Reasoning, Method=FTRL-Reinforce++2025.08 | 27.09 | — | — | — | |
| Qwen2.5-14B FTRL-Reinforce++Family=Qwen2.5-14B, Method=FTRL-Reinforce++2025.08 | 26.83 | — | — | — | |
| Qwen2.5-14B FTRL-GRPOFamily=Qwen2.5-14B, Method=FTRL-GRPO2025.08 | 25.43 | — | — | — | |
| Qwen3-14B Non-Reasoning FTRL-GRPOFamily=Qwen3-14B, Version=Non-Reasoning, Method=FTRL-GRPO2025.08 | 24.26 | — | — | — | |
| Qwen3-8B Non-Reasoning FTRL-GRPOFamily=Qwen3-8B, Version=Non-Reasoning, Method=FTRL-GRPO2025.08 | 23.35 | — | — | — | |
| Qwen3-8B Non-Reasoning FTRL-Reinforce++Family=Qwen3-8B, Version=Non-Reasoning, Method=FTRL-Reinforce++2025.08 | 21.96 | — | — | — | |
| Qwen2.5-32B InstructFamily=Qwen2.5-32B, Version=Instruct2025.08 | 21.91 | — | — | — | |
| Qwen3-14B ReasoningFamily=Qwen3-14B, Version=Reasoning2025.08 | 18.87 | — | — | — | |
| Qwen3-14B Non-Reasoning FTRL-Reinforce++Family=Qwen3-14B, Version=Non-Reasoning, Method=FTRL-Reinforce++2025.08 | 17.61 | — | — | — | |
| Qwen2.5-14B InstructFamily=Qwen2.5-14B, Version=Instruct2025.08 | 16.74 | — | — | — | |
| Qwen3-8B ReasoningFamily=Qwen3-8B, Version=Reasoning2025.08 | 16.43 | — | — | — | |
| GPT 3.5-TurboVersion=3.5-Turbo2025.08 | 15.13 | — | — | — | |
| Qwen3-14B Non-ReasoningFamily=Qwen3-14B, Version=Non-Reasoning2025.08 | 13.74 | — | — | — | |
| Qwen2.5-7B FTRL-Reinforce++Family=Qwen2.5-7B, Method=FTRL-Reinforce++2025.08 | 11.91 | — | — | — | |
| Qwen3-8B Non-ReasoningFamily=Qwen3-8B, Version=Non-Reasoning2025.08 | 10.13 | — | — | — | |
| Qwen2.5-7B FTRL-GRPOFamily=Qwen2.5-7B, Method=FTRL-GRPO2025.08 | 6.91 | — | — | — | |
| Qwen2.5-7B InstructFamily=Qwen2.5-7B, Version=Instruct2025.08 | 5.91 | — | — | — |