Loading the SOTA2 catalog…
DiffVC: A Non-autoregressive Framework Based on Diffusion Model for Video Captioning · SOTA2 Research