ResearchBenchmarksMany-to-one Text Generation ((I+A) → T) on Triplet (text, image, audio)Follow27.83CLIP Score (I->T)FlowBind23.888424.911725.93526.9583Dec 17, 2025Evaluation ResultsMethodMethodLinksCLIP Score (I->T)CLAP Score (A->T)FlowBind2025.1227.8335.21OmniFlow2025.1226.3836.07CoDi2025.1224.0420.66