Loading the SOTA2 catalog…
VLCap: Vision-Language with Contrastive Learning for Coherent Video Paragraph Captioning · SOTA2 Research