ResearchTasksAudio-to-(Text+Image) GenerationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast Updatedtriplet (text, image, audio)FlowBind36.79CLAP Score (A->T)3Feb 26, 2026