Loading the SOTA2 catalog…
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning · SOTA2 Research