Loading the SOTA2 catalog…
Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes · SOTA2 Research