ResearchTasksSpeech-to-Image RetrievalFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedPlaces audio caption dataset 1,000 image/caption (held out)MISA0.271R@114Feb 26, 2026