Text Generation
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
24.25Lexical Divergence
7
Feb 26, 2026
25.91Lexical Divergence
7
Feb 26, 2026
23.07Lexical Divergence
7
Feb 26, 2026
3.2Total Training Time (h)
6
Jun 30, 2026
5Coherent Emission Rate
6
May 15, 2026
3.2ROUGE-1
6
May 13, 2026
42.24Perplexity
6
May 12, 2026
43.27Perplexity
6
May 12, 2026
33.51Perplexity
6
May 12, 2026
31.04Perplexity
6
May 12, 2026
41.48BLEU-3
6
Apr 15, 2026
45.79BLEU-1
6
Mar 10, 2026
49.2BLEU
6
Feb 26, 2026
3.973MedHelm Jury Score
6
Feb 26, 2026
4.414MedHelm LLM-jury score
6
Feb 26, 2026
3.882MedHelm LLM-Jury Score
6
Feb 26, 2026
25.3BLEU
6
Feb 26, 2026
69.23Sen. ACC
6
Feb 26, 2026
72.39Sentence Accuracy
6
Feb 26, 2026
0.84BERTscore
6
Feb 26, 2026
0.934BERTScore
6
Feb 26, 2026
0.929BERTScore
6
Feb 26, 2026
90.5BLEU-2
6
Feb 26, 2026
6.14Oracle NLL
6
Feb 26, 2026
5.67Oracle NLL
6
Feb 26, 2026