ResearchTasksGeneral Language CapabilitiesFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedOpen LLM Leaderboard lm-eval-harness (test)TDPO83.29HellaSwag Accuracy14May 13, 2026MMLU, GSM8K, GPQA, HumanEval, TruthfulQA, IFEval AggregateGRPO71.2Average Score10Mar 4, 2026