Loading the SOTA2 catalog…
OpenViVQA: Task, Dataset, and Multimodal Fusion Models for Visual Question Answering in Vietnamese · SOTA2 Research