Loading the SOTA2 catalog…
VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning · SOTA2 Research