Loading the SOTA2 catalog…
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks · SOTA2 Research