Dialogue Response Generation
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
35.8B-4 Score
38
Apr 10, 2026
33.5B-4
38
Apr 10, 2026
53.3BLEU-1
20
Feb 26, 2026
98.5Und
16
Feb 26, 2026
45.84F1
14
May 21, 2026
47.43F1
14
May 21, 2026
5.47Perplexity
14
Feb 26, 2026
22.89F1
12
Feb 26, 2026
24.18BLEU-1 Score
10
Jun 12, 2026
92Per-response Accuracy
9
Feb 26, 2026
100Per-response Accuracy
9
Feb 26, 2026
96.7Accuracy (Per Response)
9
Feb 26, 2026
100Accuracy (Per-response)
9
Feb 26, 2026
1Per-response Accuracy
9
Feb 26, 2026
99.6Per-response Accuracy
9
Feb 26, 2026
100Per-response Accuracy
9
Feb 26, 2026
96.3Accuracy (Per-response)
9
Feb 26, 2026
100Per-response accuracy
9
Feb 26, 2026
100Per-response Accuracy
9
Feb 26, 2026
29.1BLEU-4
8
Apr 10, 2026
4.56Perplexity
7
Feb 26, 2026
37.2Adversarial Success
7
Feb 26, 2026
4.7BLEU Score (Month 9)
6
Feb 27, 2026
7.26BLEU (Month 9)
6
Feb 27, 2026
66.1Appropriateness
6
Feb 26, 2026