Loading the SOTA2 catalog…
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering · SOTA2 Research