Loading the SOTA2 catalog…
HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification · SOTA2 Research