Interactive tool-use on τ-bench retail domain
28.7Action RewardGBC (Qwen-3-32B)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GBC (Qwen-3-32B)Architecture=Multi-Agent, Optimization Formula=Max of L1 Norm2026.06 | 28.7 | 79.1 | 24.3 | |
| Qwen-3-32BArchitecture=Single-Agent2026.06 | 27.8 | 71.3 | 22.6 | |
| GBC (Qwen-3-32B)Architecture=Multi-Agent, Optimization Formula=Mean of Product with Input2026.06 | 27 | 78.3 | 24.3 | |
| GBC (Qwen-3-32B)Architecture=Multi-Agent, Optimization Formula=Max of Product with Input2026.06 | 20 | 77.4 | 17.4 | |
| GBC (Qwen-3-32B)Architecture=Multi-Agent, Optimization Formula=Mean of L1 Norm2026.06 | 17.4 | 73 | 13.9 | |
| Qwen-3-32BArchitecture=Multi-Agent, Optimization Status=Before Optimization2026.06 | 13.9 | 78.3 | 13 | |
| GBC (Llama-3.3-70B-It)Architecture=Multi-Agent, Optimization Formula=Max of L1 Norm2026.06 | 13.9 | 73 | 9.6 | |
| GBC (Llama-3.3-70B-It)Architecture=Multi-Agent, Optimization Formula=Max of Product with Input2026.06 | 13.9 | 71.3 | 9.6 | |
| GBC (Llama-3.3-70B-It)Architecture=Multi-Agent, Optimization Formula=Mean of Product with Input2026.06 | 12.2 | 72.2 | 8.7 | |
| GBC (Llama-3.3-70B-It)Architecture=Multi-Agent, Optimization Formula=Mean of L1 Norm2026.06 | 11.3 | 70.4 | 8.7 | |
| Llama-3.3-70B-ItArchitecture=Single-Agent2026.06 | 9.6 | 62.6 | 7 | |
| Llama-3.3-70B-ItArchitecture=Multi-Agent, Optimization Status=Before Optimization2026.06 | 9.6 | 70.4 | 6.1 |