Loading the SOTA2 catalog…
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering · SOTA2 Research