Loading the SOTA2 catalog…
Efficient Large Multi-modal Models via Visual Context Compression · SOTA2 Research