Loading the SOTA2 catalog…
Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding · SOTA2 Research