Loading the SOTA2 catalog…
PaCa-ViT: Learning Patch-to-Cluster Attention in Vision Transformers · SOTA2 Research