Loading the SOTA2 catalog…
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos · SOTA2 Research