ResearchTasksLong-context language tasks (MC, QA, Sum)FollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast Updated∞BenchRAG78.6MC Accuracy13Feb 26, 2026