Loading the SOTA2 catalog…
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs · SOTA2 Research