Loading the SOTA2 catalog…
EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning · SOTA2 Research