ResearchTasksMany-to-one Text Generation ((I+A) → T)FollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedTriplet (text, image, audio)FlowBind27.83CLIP Score (I->T)3Feb 26, 2026