Agentic Tool-use on tau^2 Bench (GPT-4.1 Simulator Setting)
0.775Retail ScoreREACT(GPT-5)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| REACT(GPT-5)Backbone=GPT-5, User Simulator=GPT-4.1, Maximum Step=1002026.02 | 0.775 | 0.975 | 0.517 | 0.803 | |
| REACT(GPT-4.1)Backbone=GPT-4.1, User Simulator=GPT-4.1, Maximum Step=1002026.02 | 0.667 | 0.5 | 0.417 | 0.55 | |
| SRTraining Category=Learning from experts/strong LLMs, User Simulator=GPT-4.1, Maximum Step=1002026.02 | 0.525 | 0.458 | 0.433 | 0.48 | |
| Imitation LearningTraining Category=Learning from experts/strong LLMs, User Simulator=GPT-4.1, Maximum Step=1002026.02 | 0.492 | 0.5 | 0.333 | 0.463 | |
| RWML + Policy RLTraining Category=Self-supervised + Policy RL, User Simulator=GPT-4.1, Maximum Step=1002026.02 | 0.483 | 0.417 | 0.5 | 0.46 | |
| RWMLTraining Category=Self-supervised, User Simulator=GPT-4.1, Maximum Step=1002026.02 | 0.433 | 0.475 | 0.45 | 0.453 | |
| REACT(Qwen3-8B)Backbone=Qwen3-8B, User Simulator=GPT-4.1, Maximum Step=1002026.02 | 0.425 | 0.3 | 0.233 | 0.337 | |
| WM SFTTraining Category=Self-supervised, User Simulator=GPT-4.1, Maximum Step=1002026.02 | 0.408 | 0.308 | 0.3 | 0.347 | |
| Policy RLTraining Category=Learning from task success reward, User Simulator=GPT-4.1, Maximum Step=1002026.02 | 0.342 | 0.45 | 0.317 | 0.38 |