ResearchBenchmarksImage-to-(Text+Audio) Generation on triplet dataset text, image, audioFollow27.98CLIP Score (I->T)FlowBind25.6426.247526.85527.4625Dec 17, 2025Evaluation ResultsMethodMethodLinksCLIP Score (I->T)AIS Score (I->A)FlowBind2025.1227.9874.34OmniFlow2025.1226.3663.99CoDi2025.1225.7358.65