Loading the SOTA2 catalog…
ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders · SOTA2 Research