Loading the SOTA2 catalog…
VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks · SOTA2 Research