Loading the SOTA2 catalog…
Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment · SOTA2 Research