Loading the SOTA2 catalog…
CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment · SOTA2 Research