ResearchDatasetsStandard LLM BenchmarksFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsReasoning and Question AnsweringStandard LLM Benchmarks (BoolQ, RTE, HellaSWAG, ARC, OpenBookQA, PiQA)67.24Avg Accuracy15
Reasoning and Question AnsweringStandard LLM Benchmarks (BoolQ, RTE, HellaSWAG, ARC, OpenBookQA, PiQA)67.24Avg Accuracy15