Loading the SOTA2 catalog…
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models · SOTA2 Research