ResearchTasksAudio-to-Text GenerationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedImage-Text-AudioFlowBind54.41CIDEr9Jun 16, 2026one-to-one evaluation benchmarksOmniFlow45.08CLAP Score5Feb 26, 2026One-to-one evaluation benchmarks Audio-to-TextFlowBind55.11CIDEr5Feb 26, 2026