Loading the SOTA2 catalog…
M$^{3}$3D: Learning 3D priors using Multi-Modal Masked Autoencoders for 2D image and video understanding · SOTA2 Research