ResearchTasksInstruction following and reasoningFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedLow-resource languages evaluation suite (am, arz, ars, as, ast, az, ba, bn, bo, ceb, cv, cy, fo, ga, gd, gl, gn, ha, ht, ig, jv, kmr, sdh, ky, lb, lo, lus, mg, mi, mn, mt, ny, oc, pap, ps, rn, rw, sd, si, sm, sn, st, su, sw, te, tg, ti, tk, tt, ug, xh, yi, yo, zu)Kakugo5Wins54Feb 26, 2026Average of 9 tasks (DollyEval, VicunaEval, GSM8K, MATH, AIME2024, HumanEval, MBPP, LiveCodeBench, GPQA-D)IOA31.19Average Performance9Feb 27, 2026Chat and Instruction-following Suite IFEval, AE2, MTB, GSM8KS2FT (Down)0.695IFEval5Feb 26, 2026
Low-resource languages evaluation suite (am, arz, ars, as, ast, az, ba, bn, bo, ceb, cv, cy, fo, ga, gd, gl, gn, ha, ht, ig, jv, kmr, sdh, ky, lb, lo, lus, mg, mi, mn, mt, ny, oc, pap, ps, rn, rw, sd, si, sm, sn, st, su, sw, te, tg, ti, tk, tt, ug, xh, yi, yo, zu)Kakugo5Wins54Feb 26, 2026
Average of 9 tasks (DollyEval, VicunaEval, GSM8K, MATH, AIME2024, HumanEval, MBPP, LiveCodeBench, GPQA-D)IOA31.19Average Performance9Feb 27, 2026