Dialog Generation
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
0.832BLEU-1
12
Feb 26, 2026
47.4Accuracy (Response)
10
Feb 26, 2026
6.36Perplexity (PPL)
7
May 22, 2026
4.37Perplexity (PPL)
7
May 22, 2026
4.73Appropriateness
3
Feb 26, 2026
4.73Appropriateness
3
Feb 26, 2026