ResearchDatasetsLong-Context Evaluation SuiteFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsLong-context language modelingLong-Context Evaluation Suite MRCR v2, GraphWalks, LongBench v2, RULER, AA-LCR78.7Average Score5
Long-context language modelingLong-Context Evaluation Suite MRCR v2, GraphWalks, LongBench v2, RULER, AA-LCR78.7Average Score5