End-to-End Task-Oriented Dialogue on MultiWOZ 2.1 (test)
22.19BLEU ScoreDarwinTOD
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DarwinTODBackbone=GPT-5.12026.01 | 22.19 | 99.4 | 96.5 | 120.14 | |
| ESAinsTOD (GT)Base Model=Llama 2 7B, ground-truth system actions=true2026.03 | 21.92 | 87.6 | 81.1 | 106.27 | |
| DarwinTODBackbone=GPT-42026.01 | 21.46 | 99.76 | 96.38 | 119.53 | |
| ESAinsTODBase Model=Llama 2 7B2026.03 | 21.41 | 94.4 | 87.5 | 112.38 | |
| Llama 2 7B (FT)fine-tuned exclusively=true2026.03 | 21.33 | 92.3 | 84.6 | 109.78 | |
| DarwinTODBackbone=Qwen3-14B2026.01 | 20.88 | 99.84 | 94.92 | 118.26 | |
| DarwinTODBackbone=Qwen2.5-14B2026.01 | 20.67 | 99.18 | 92.64 | 116.58 | |
| DarwinTODBackbone=Qwen3-8B2026.01 | 20.33 | 99.62 | 94.18 | 117.23 | |
| DarwinTODBackbone=Llama3.1-8B2026.01 | 20.32 | 97.97 | 91.51 | 115.06 | |
| GALAXYbelief_states=oracle2021.11 | 20.29 | 94.8 | 86.2 | 110.79 | |
| DarwinTODBackbone=Qwen2.5-7B2026.01 | 20.18 | 98.92 | 91.85 | 115.57 | |
| MTTODsetting=end-to-end, reproducibility=reproduced results2023.05 | 20.15 | 90.4 | 81.7 | 106.2 | |
| DarwinTODBackbone=Qwen3-4B2026.01 | 20.04 | 98.76 | 91.62 | 115.23 | |
| GALAXYPre-training=true2021.11 | 20.01 | 95.3 | 86.2 | 110.76 | |
| DarwinTODBackbone=Llama3-8B2026.01 | 19.96 | 98.73 | 91.42 | 115.04 | |
| AgentTODBackbone=Llama3-8B2026.01 | 19.91 | 96.89 | 89.38 | 113.04 | |
| SPACEfine-tuned individually=true2026.03 | 19.91 | 93 | 84.1 | 108.46 | |
| MTTOD2026.03 | 19.68 | 90.99 | 82.08 | 106.22 | |
| MarCobelief_states=oracle2021.11 | 19.54 | 92.5 | 77.8 | 104.69 | |
| PPTOD base2021.09 | 19.17 | 87.09 | 79.08 | 102.26 | |
| PPTOD2021.11 | 19.17 | 87.09 | 79.08 | 102.26 | |
| PPTODsetting=end-to-end2023.05 | 19.17 | 87.09 | 79.08 | 102.26 | |
| PPTODfine-tuned individually=true2026.03 | 19.17 | 87.09 | 79.08 | 102.26 | |
| MDTODsetting=end-to-end2023.05 | 19.03 | 92.7 | 84.6 | 107.68 | |
| HDNObelief_states=oracle2021.11 | 18.97 | 92.8 | 83 | 106.87 | |
| PPTOD small2021.09 | 18.59 | 88.89 | 76.98 | 101.52 | |
| GALAXY(w/o pre-train)belief_states=oracle, initialization=original weights of UniLM2021.11 | 18.58 | 93.7 | 83.3 | 107.08 | |
| DarwinTODBackbone=Qwen2.5-3B2026.01 | 18.42 | 92.15 | 83.84 | 106.42 | |
| GALAXYInitialization strategy=original weights of UniLM, Pre-training=false2021.11 | 18.32 | 93.5 | 81.7 | 105.92 | |
| GALAXYsetting=end-to-end, pre-training=without2023.05 | 18.32 | 93.5 | 81.7 | 105.92 | |
| LABES-S2S2021.09 | 18.13 | 75.07 | 67.06 | 89.19 | |
| DAMDBelief State=Generated2021.05 | 18 | 72.4 | 57.7 | 83.05 | |
| PPTOD large2021.09 | 17.89 | 86.43 | 74.35 | 98.28 | |
| GPT2 + NeuralWOZBelief State=Oracle2021.05 | 17.69 | 78.1 | 67.6 | 90.54 | |
| GPT2 + NeuralWOZBelief State=Generated2021.05 | 17.46 | 75.1 | 64.6 | 87.31 | |
| GPT2 (baseline)Belief State=Generated2021.05 | 17.38 | 74.6 | 64.4 | 86.88 | |
| DAMDBelief State=Oracle2021.05 | 17.3 | 80.3 | 65.1 | 90 | |
| GPT2 (baseline)Belief State=Oracle2021.05 | 17.27 | 77.1 | 67.8 | 89.72 | |
| UBARbelief_states=oracle2021.11 | 16.7 | 92.7 | 81 | 103.55 | |
| UBAR2021.11 | 16.5 | 95.7 | 81.8 | 105.25 | |
| UBAROracle dialogue states history=true2021.09 | 16.48 | 86.2 | 70.32 | 94.74 | |
| UBARsetting=end-to-end, result_source=Su et al., 20222023.05 | 16.48 | 86.2 | 70.32 | 94.74 | |
| UBARoracle dialogue states=true2026.03 | 16.48 | 86.2 | 70.32 | 94.74 | |
| SimpleTODBelief State=Oracle2021.05 | 16.22 | 85.1 | 73.5 | 95.52 | |
| SimpleTODbelief_states=oracle2021.11 | 16.22 | 85.1 | 73.5 | 95.52 | |
| GPT2Belief State=Oracle, Reference=Mohapatra et al., 20202021.05 | 15.95 | 72.8 | 63.7 | 84.2 | |
| GPT2Belief State=Generated, Reference=Mohapatra et al., 20202021.05 | 15.94 | 66.2 | 55.4 | 76.74 | |
| DoTS2021.11 | 15.9 | 86.65 | 74.18 | 96.32 | |
| DOTSsetting=end-to-end2023.05 | 15.9 | 86.65 | 74.18 | 96.32 | |
| SimpleTOD2021.09 | 15.23 | 85 | 70.5 | 92.98 | |
| SimpleTOD2021.11 | 15.23 | 85 | 70.5 | 92.98 | |
| SimpleTODsetting=end-to-end2023.05 | 15.23 | 85 | 70.5 | 92.98 | |
| SimpleTOD2026.03 | 15.23 | 85 | 70.5 | 92.98 | |
| GPT2 + SimulatedChatBelief State=Oracle2021.05 | 15.06 | 80.4 | 62.2 | 86.36 | |
| SimpleTODBelief State=Generated2021.05 | 14.99 | 83.4 | 67.1 | 90.24 | |
| GPT2 + SimulatedChatBelief State=Generated2021.05 | 14.62 | 72.5 | 53.7 | 77.72 | |
| LAVAbelief_states=oracle2021.11 | 14.02 | 96.39 | 83.57 | 104 |