Loading the SOTA2 catalog…
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders · SOTA2 Research