Dialogue Evaluation
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
12.4Perplexity
4
Feb 26, 2026
10.2Perplexity
4
Feb 26, 2026
97.7USR RET
4
Feb 26, 2026
0.998USR RET
4
Feb 26, 2026
75Engagingness
4
Feb 26, 2026
0.173Spearman's rho
3
Apr 13, 2026
0.04SentBLEU
3
Feb 26, 2026
0.17SentBLEU
3
Feb 26, 2026
0.91Pearson
3
Feb 26, 2026
83Divergence
2
Jun 2, 2026
86Divergence
2
Jun 2, 2026
0.806Naturalness (Pearson r)
2
Feb 26, 2026