Loading the SOTA2 catalog…
PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining · SOTA2 Research