Loading the SOTA2 catalog…
Voila-A: Aligning Vision-Language Models with User's Gaze Attention · SOTA2 Research