User Simulator Goal Alignment on τ-Bench Retail (test)
94.5User Profile Success RateGemma-2-27B-It
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemma-2-27B-ItStrategy=Inference-Time Steering2025.07 | 94.5 | 56.6 | 97.1 | 99.5 | 89.8 | 87.5 | — | — | — | — | |
| Llama-3.3-70B-ItStrategy=Inference-Time Steering2025.07 | 92.7 | 59.2 | 96 | 98.4 | 100 | 89.3 | — | — | — | — | |
| Llama-3.3-70B-ItStrategy=Prompt-Based2025.07 | 91.7 | 46.1 | 99.6 | 100 | 100 | 87.5 | — | — | — | — | |
| Gemma-2-27B-ItStrategy=Prompt-Based2025.07 | 91 | 39.5 | 97.8 | 99.5 | 96.3 | 84.8 | — | — | — | — | |
| Qwen-2.5-72B-ItStrategy=Prompt-Based2025.07 | 88.6 | 40.8 | 98.2 | 97.9 | 91.7 | 83.4 | — | — | — | — | |
| Qwen-2.5-72B-ItStrategy=Inference-Time Steering2025.07 | 85.8 | 48.7 | 97.1 | 98.4 | 98.1 | 85.6 | — | — | — | — | |
| Qwen-2.5-7B-ItStrategy=Cold-Start SFT2025.07 | 85.8 | 43.4 | 92.4 | 94.7 | 82.4 | 79.7 | — | — | — | — | |
| Qwen-2.5-7B-ItStrategy=GRPO with UGST Rewards2025.07 | 85.5 | 38.2 | 97.1 | 97.9 | 97.2 | 83.2 | — | — | — | — | |
| Llama-3.1-8B-ItStrategy=Prompt-Based2025.07 | 84.8 | 36.8 | 97.1 | 98.9 | 96.3 | 82.8 | — | — | — | — | |
| Llama-3.1-8B-ItStrategy=Inference-Time Steering2025.07 | 84.8 | 52.6 | 95.3 | 98.9 | 97.2 | 85.8 | — | — | — | — | |
| Llama-3.1-8B-ItStrategy=Cold-Start SFT2025.07 | 84.8 | 47.3 | 95.5 | 97.8 | 94.1 | 83.9 | — | — | — | — | |
| Llama-3.1-8B-ItStrategy=GRPO with UGST Rewards2025.07 | 84.8 | 47.4 | 98.9 | 99.5 | 98.1 | 85.7 | — | — | — | — | |
| HumansUser Simulator=Qwen3-Next-80B-A3B-Instruct2026.05 | 84.3 | 97.6 | 87.2 | 82.2 | — | — | 95.3 | 61.4 | 78.3 | 87.8 | |
| Qwen-2.5-7B-ItStrategy=Prompt-Based2025.07 | 82 | 40.8 | 96 | 99.5 | 91.7 | 82 | — | — | — | — | |
| Qwen-2.5-7B-ItStrategy=Inference-Time Steering2025.07 | 73 | 47.4 | 94.9 | 97.9 | 93.5 | 81.4 | — | — | — | — | |
| PPol: EvolvedUser Simulator=Qwen3-Next-80B-A3B-Instruct2026.05 | 69.6 | 89.8 | 70.4 | 76.4 | — | — | 78.4 | 60.2 | 69.3 | 76.5 | |
| PPol: InitialUser Simulator=Qwen3-Next-80B-A3B-Instruct2026.05 | 31.4 | 58.7 | 38.7 | 30.3 | — | — | 35.6 | 1.7 | 18.6 | 39.8 | |
| DP PersonasUser Simulator=Qwen3-Next-80B-A3B-Instruct2026.05 | 29.5 | 56.5 | 37.8 | 25.2 | — | — | 29.1 | 10 | 19.6 | 37.2 | |
| Base-simulatorUser Simulator=Qwen3-Next-80B-A3B-Instruct2026.05 | 24 | 57.5 | 35.7 | 23.3 | — | — | 10.7 | 4.6 | 7.7 | 35.1 |