ResearchTasksLanguage Understanding and Code GenerationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedLlama 1B Evaluation Suite (ARC, HellaSwag, MMLU, TruthfulQA, WinoGrande, Humaneval) 3.2QLoRA w/ TOKENTUNE (Random)39.33ARC6Feb 26, 2026
Llama 1B Evaluation Suite (ARC, HellaSwag, MMLU, TruthfulQA, WinoGrande, Humaneval) 3.2QLoRA w/ TOKENTUNE (Random)39.33ARC6Feb 26, 2026