Loading the SOTA2 catalog…
Unified Vision-Language Pre-Training for Image Captioning and VQA · SOTA2 Research