ResearchTasksText-to-(Image+Audio) GenerationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast Updatedtriplet (text, image, audio)CoDi26.61CLIP Score (T->I)3Feb 26, 2026