Loading the SOTA2 catalog…
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model · SOTA2 Research