Loading the SOTA2 catalog…
ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages · SOTA2 Research