NarrativeQA
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
NarrativeQA Peter Rabbit
16Exact Match
4
NarrativeQA
3.212R@1
4
NarrativeQA (dev)
31.6ROUGE-L
4
NarrativeQA (test)
58.8ROUGE-L
4
NarrativeQA Story Summaries (test)
42BLEU-1
4
NarrativeQA LongBench (test)
11.48String Match
3
NarrativeQA SCROLLS (test)
11.63F1 Score
3
NarrativeQA
88.5Score (avg@3)
3
NarrativeQA Robinson Crusoe
40Exact Match
3
NarrativeQA The Phantom of the Opera 25k nodes, 56k edges
34Exact Match (EM)
3
NarrativeQA
97.71Latency (s)
3
NarrativeQA (dev)
58.1ROUGE-L
3
NarrativeQA (dev)
25.4EM
2