ResearchDatasetsLAM Evaluation BenchmarkFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsBenchmark Subset SelectionLAM Evaluation Benchmark 40 tasks0.977Pearson Correlation60