Task-oriented Dialogue on MultiWOZ 89 scenarios
92.7Collaboration SRGPT-4.1-mini
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4.1-miniReasoning step limit=30, Number of trials=42025.09 | 92.7 | 89.3 | 89.3 | 90.7 | 88.2 | 100 | 96.3 | 96.3 | 97.8 | 95.1 | |
| Qwen3-235b-a22bReasoning step limit=30, Number of trials=42025.09 | 77.8 | 62.4 | 57.3 | 69.4 | 69.9 | 100 | 80.2 | 73.7 | 89.2 | 89.8 | |
| Llama-3.1-70b-instructReasoning step limit=30, Number of trials=42025.09 | 62.6 | 54.8 | 49.4 | 47.5 | 48.6 | 100 | 87.5 | 78.9 | 75.9 | 77.6 | |
| Qwen3-30b-a3bReasoning step limit=30, Number of trials=42025.09 | 48.3 | 47.2 | 27.2 | 41 | 26.1 | 100 | 97.7 | 56.3 | 84.9 | 54 | |
| GPT-4.1-nanoReasoning step limit=30, Number of trials=42025.09 | 23.6 | 16.9 | 9.8 | 26.7 | 14.7 | 100 | 71.6 | 41.5 | 113.1 | 62.3 |