ResearchTasksZero-shot NLP EvaluationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedAggregateVERSATILEFFN60.47Average Accuracy18Feb 26, 2026NLP Downstream Benchmarks (ARC, BoolQ, HellaSwag, LAMBADA, MMLU)Muon-320.354ARC-C15Feb 26, 2026