RULER
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
RULER Multi-Key (test)
99.5Accuracy (4K Context)
8
RULER S-NIAH-3 2K context
100Recall Accuracy
8
RULER S-NIAH-3 (1K context)
100Recall Accuracy
8
RULER S-NIAH-3 (0.5K context)
100Recall Accuracy
8
RULER S-NIAH-2 2K
100Recall Accuracy
8
RULER S-NIAH-2 (1K context)
100Recall Accuracy
8
RULER S-NIAH-2 (0.5K context)
100Recall Accuracy
8
RULER S-NIAH-1 4K context
100Recall Accuracy
8
RULER S-NIAH-1 2K
100Recall Accuracy
8
RULER S-NIAH-1 1K context
100Recall Accuracy
8
RULER S-NIAH-1 (0.5K)
100Recall Accuracy
8
RULER Sequence length = 64k
100S-NIAH Score (Component 1)
8
RULER (test)
100S1 Score
8
RULER QA-8k (test)
512Token Count
8
RULER QA-16k (test)
512Token Count
8
Ruler S-NIAH
100Accuracy
8
Ruler
100Retrieval Accuracy @ 1024 Context
8
RULER 4K-32K context
95.73Accuracy
8
RULER (test)
2,800Multi-Query Success Rate
8
RULER
3.43Mean Throughput
8
RULER 4k context length (test)
25.36MK
7
RULER 32K
100S-N Score
7
RULER (test)
96.6Accuracy (4k Context)
7
RULER 8k context
100Score 1 (S1)
7
RULER 4k context
100S1 Score
7