Loading the SOTA2 catalog…
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference · SOTA2 Research