ResearchTasksContextual ExtractionFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedBalanced evaluation dataset NQ-Open, TriviaQA, HotpotQA 1,000 samples per corpus (test)GPT-4o-mini62Accuracy (Zero-Shot)4Jun 24, 2026
Balanced evaluation dataset NQ-Open, TriviaQA, HotpotQA 1,000 samples per corpus (test)GPT-4o-mini62Accuracy (Zero-Shot)4Jun 24, 2026