Dialogue Generation
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
4.21Correctness
3
Feb 26, 2026
100AU
2
Feb 26, 2026
0.749Distinct-2
2
Feb 26, 2026
10.04BLEU-4
2
Feb 26, 2026
0.86Embodiment Score
2
Feb 26, 2026
44.1ROUGE-2
2
Feb 26, 2026
19Fluency (Win)
2
Feb 26, 2026
90Proportion
1
May 20, 2026
49.51Relevance Win %
1
Feb 26, 2026
41.34Relevance Win Rate
1
Feb 26, 2026
—
—Primary metric
0
Feb 26, 2026
—
—Attribute Relevancy
0
Feb 26, 2026
—
—Primary metric
0
Feb 26, 2026