Loading the SOTA2 catalog…
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language · SOTA2 Research