ResearchTasks(I+A)→T Cross-modal AlignmentFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedImage-Text-AudioUnifiedIO2-L29.97CLIP Score16Jun 16, 2026