HELMET
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
HELMET
55.2Average Score
39
HELMET
78.4Accuracy
37
HELMET
0Average Sparsity
28
HELMET
247Summarization Score
27
HELMET
33.1Score
25
HELMET 2025
61.44Accuracy (8K Context)
16
HELMET multi_lexsum
94.9Utilization
15
HELMET
68.5Accuracy
15
HELMET held-out eval
57.61Accuracy (8K Context)
13
HELMET RAG subset
81.1HotpotQA Accuracy
8
Helmet
67.6Accuracy
6
HELMET shorter context lengths ≤32K
59.34Score (8K Context)
4
HELMET Holistic understanding
46.51HELMET Holistic Understanding (64K Context)
4
HELMET 128K
46.53Overall Score
2