Loading the SOTA2 catalog…
Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning · SOTA2 Research