Loading the SOTA2 catalog…
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models · SOTA2 Research