Loading the SOTA2 catalog…
VILA: Learning Image Aesthetics from User Comments with Vision-Language Pretraining · SOTA2 Research