Loading the SOTA2 catalog…
LAVT: Language-Aware Vision Transformer for Referring Image Segmentation · SOTA2 Research