Loading the SOTA2 catalog…
Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models · SOTA2 Research