Loading the SOTA2 catalog…
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding · SOTA2 Research