Loading the SOTA2 catalog…
How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Heads · SOTA2 Research