Loading the SOTA2 catalog…
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding · SOTA2 Research