ResearchTasksLanguage Modeling and Zero-shot ReasoningFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedStandard LLM Evaluation Suite ARC-e, PIQA, Hellaswag, OpenBookQA, Winogrande, MMLU, BoolQBHyT3.254PT Eval Loss5Feb 26, 2026
Standard LLM Evaluation Suite ARC-e, PIQA, Hellaswag, OpenBookQA, Winogrande, MMLU, BoolQBHyT3.254PT Eval Loss5Feb 26, 2026