ResearchDatasetsMMLU, ARC-Challenge, and CommonsenseQAFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsGeneral Language ModelingMMLU, ARC-Challenge, and CommonsenseQA Aggregate64.77Average Score24