Loading the SOTA2 catalog…
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling · SOTA2 Research