Loading the SOTA2 catalog…
Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos · SOTA2 Research