Loading the SOTA2 catalog…
Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning · SOTA2 Research