Loading the SOTA2 catalog…
E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning · SOTA2 Research