ResearchTasksMany-to-one Audio Generation ((T+I) → A)FollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedTriplet text, image, audioFlowBind28.13CLAP Score (T->A)3Feb 26, 2026