Loading the SOTA2 catalog…
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models · SOTA2 Research