ResearchTasksZero-shot Language Modeling and Commonsense ReasoningFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedStandard short-context evaluation suite (LAMBADA, ARC, HellaSwag, PIQA, WinoGrande)OVQ/SWA57.3LAMBADA PPL Score18May 12, 2026
Standard short-context evaluation suite (LAMBADA, ARC, HellaSwag, PIQA, WinoGrande)OVQ/SWA57.3LAMBADA PPL Score18May 12, 2026