ResearchTasksReasoning and Question AnsweringFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedStandard LLM Benchmarks (BoolQ, RTE, HellaSWAG, ARC, OpenBookQA, PiQA)Before finetune67.24Avg Accuracy15Feb 26, 2026HLE (reference)Theoria185Problems Evaluated1Jul 2, 2026
Standard LLM Benchmarks (BoolQ, RTE, HellaSWAG, ARC, OpenBookQA, PiQA)Before finetune67.24Avg Accuracy15Feb 26, 2026