ResearchTasksImage-to-(Text+Audio) GenerationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast Updatedtriplet dataset text, image, audioFlowBind27.98CLIP Score (I->T)3Feb 26, 2026