ResearchDatasetsVisSpeechFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsAudio-Visual Automatic Speech RecognitionVisSpeech zero-shot16.6WER6