Loading the SOTA2 catalog…
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions · SOTA2 Research