Loading the SOTA2 catalog…
Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? · SOTA2 Research