Loading the SOTA2 catalog…
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval · SOTA2 Research