Loading the SOTA2 catalog…
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding · SOTA2 Research