ResearchDatasetsDH-FaceVidFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsJoint text-to-audio-video generationDH-FaceVid-1K Chinese (test)0.148CER6