ResearchDatasetsCombined Image and Video BenchmarksFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsGeneral Multi-modal UnderstandingCombined Image and Video Benchmarks100Average Accuracy16