ResearchBenchmarksSynthetic Dialogue Generation Evaluation on MTS-dialog (test)Follow52.7ROUGE-1 (Source-Hypothesis)GPT434.593639.294343.99548.6957Oct 24, 2023Evaluation ResultsMethodMethodLinksROUGE-1 (Source-Hypothesis)ROUGE-1ROUGE-2ROUGE-LsumConcept PrecisionConcept RecallConcept F1ROUGE-2 (Source-Hypothesis)ROUGE-Lsum (Source-Hypothesis)sBLEU (All)sBLEU (Physician)sBLEU (Patient)GPT42023.1052.753.2920.250.8171.4645.6955.1725.749.630.0190.0090.019ChatGPT2023.1043.7348.5616.7446.3667.5435.7546.2319.7240.540.0170.0060.017NoteChat2023.1037.2456.4819.7453.4148.2351.2349.6820.8336.040.0140.0070.014Human2023.1035.29——————14.3832.89———