Loading the SOTA2 catalog…
FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks · SOTA2 Research