Loading the SOTA2 catalog…
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation · SOTA2 Research