ResearchTasksGeneral Multi-task Language and Vision UnderstandingFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast Updated24-benchmark suite S3 text-friendly (test)LLM54Macro Avg Score5May 11, 202624-benchmark suite S2 visual-friendly (test)VLM + Foveation44.7Macro Score5May 11, 202624-benchmark suite All (test)Oracle-routed50.2Macro-average Score5May 11, 2026