ResearchDatasetsLLM Evaluation BenchmarksFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsLanguage Modeling AccuracyLLM Evaluation Benchmarks Zero-shot68.8Llama-2 7B Accuracy9